{
  "id": 218772,
  "title": "How important are annotations?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/218772",
  "author_name": "Gianluca Rossi",
  "post_date": "2021-02-12T04:11:10.119000",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I'm just getting started in this competition. So far, I've been able to get a 0.95 OOF CV score (GroupKFold) by just using 512x512 images and a relatively lightweight model (EfficientNet B4). I haven't had time to experiment with annotations yet.</p>\n<p>I've seen quite a few higher scores on the leaderboard. I'm wondering how important have been annotations for you? How much have they improved your score?</p>\n<p>Many thanks!</p>",
  "messages": [
    {
      "id": 1197237,
      "postDate": "2021-02-12T04:11:10.120Z",
      "content": "<p>Hi,</p>\n<p>I'm just getting started in this competition. So far, I've been able to get a 0.95 OOF CV score (GroupKFold) by just using 512x512 images and a relatively lightweight model (EfficientNet B4). I haven't had time to experiment with annotations yet.</p>\n<p>I've seen quite a few higher scores on the leaderboard. I'm wondering how important have been annotations for you? How much have they improved your score?</p>\n<p>Many thanks!</p>",
      "rawMarkdown": "Hi,\n\nI'm just getting started in this competition. So far, I've been able to get a 0.95 OOF CV score (GroupKFold) by just using 512x512 images and a relatively lightweight model (EfficientNet B4). I haven't had time to experiment with annotations yet.\n\nI've seen quite a few higher scores on the leaderboard. I'm wondering how important have been annotations for you? How much have they improved your score?\n\nMany thanks!",
      "votes": 7
    },
    {
      "id": 1201078,
      "postDate": "2021-02-15T06:36:19.633Z",
      "content": "<p>At least in my experience so far I have not seen any value coming from the annotations. Simply seems too sparse of labels and not enough relationship to be found with the things that actually matter, like the heart. Our models actually do quite well given how precise of a detection it is. </p>",
      "rawMarkdown": "At least in my experience so far I have not seen any value coming from the annotations. Simply seems too sparse of labels and not enough relationship to be found with the things that actually matter, like the heart. Our models actually do quite well given how precise of a detection it is. ",
      "votes": 1
    },
    {
      "id": 1199954,
      "postDate": "2021-02-14T09:24:02.507Z",
      "content": "<p>Make sure but to ignore <code>patient_id</code> in constructing your CV scheme, otherwise your CV results are likely to be badly misleading.</p>",
      "rawMarkdown": "Make sure but to ignore `patient_id` in constructing your CV scheme, otherwise your CV results are likely to be badly misleading."
    },
    {
      "id": 1198519,
      "postDate": "2021-02-13T05:55:16.150Z",
      "content": "<p>Thank you for looking into this <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> </p>\n<p>Yes, I’m using GroupKFold for the cross-validation strategy. </p>",
      "rawMarkdown": "Thank you for looking into this @anjum48 \n\nYes, I’m using GroupKFold for the cross-validation strategy. ",
      "replies": [
        {
          "id": 1211960,
          "postDate": "2021-02-20T17:46:16.840Z",
          "content": "<p><a href=\"https://www.kaggle.com/gianlucarossi\" target=\"_blank\">@gianlucarossi</a>  what are the annotations for ?</p>",
          "rawMarkdown": "@gianlucarossi  what are the annotations for ?\n"
        }
      ]
    },
    {
      "id": 1198336,
      "postDate": "2021-02-12T23:49:32.297Z",
      "content": "<p>0.95 is a very good CV! Did you use GroupKFold for your splits?</p>\n<p>I've just started working on this, and I'm running some tests now to see if the annotations make a difference. Will report back when I have some answers</p>",
      "rawMarkdown": "0.95 is a very good CV! Did you use GroupKFold for your splits?\n\nI've just started working on this, and I'm running some tests now to see if the annotations make a difference. Will report back when I have some answers",
      "replies": [
        {
          "id": 1199212,
          "postDate": "2021-02-13T16:04:39.227Z",
          "content": "<p>I ran 2 experiments using EfficientNet-B0 and 512x512 images, using only <code>train_annotations.csv</code> &amp; 5 folds of GroupKFold.</p>\n<p>Plain Classifer: CV = 0.859<br>\nSegmentation model with classifier head: CV = 0.881</p>\n<p>So for a weak model, adding the annotations certainly helps. I suspect that the gap will close with the larger models, but the question is how much…</p>\n<p>Edit: found an issue with my pipeline - those scores are actually for 256x256</p>",
          "rawMarkdown": "I ran 2 experiments using EfficientNet-B0 and 512x512 images, using only `train_annotations.csv` & 5 folds of GroupKFold.\n\nPlain Classifer: CV = 0.859\nSegmentation model with classifier head: CV = 0.881\n\nSo for a weak model, adding the annotations certainly helps. I suspect that the gap will close with the larger models, but the question is how much...\n\nEdit: found an issue with my pipeline - those scores are actually for 256x256"
        },
        {
          "id": 1199984,
          "postDate": "2021-02-14T10:05:05.260Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1201078,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2021-02-15T06:36:19.633000",
      "content": "<p>At least in my experience so far I have not seen any value coming from the annotations. Simply seems too sparse of labels and not enough relationship to be found with the things that actually matter, like the heart. Our models actually do quite well given how precise of a detection it is. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1199954,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-14T09:24:02.507000",
      "content": "<p>Make sure but to ignore <code>patient_id</code> in constructing your CV scheme, otherwise your CV results are likely to be badly misleading.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1198519,
      "author_name": "Gianluca Rossi",
      "author_url": "",
      "post_date": "2021-02-13T05:55:16.150000",
      "content": "<p>Thank you for looking into this <a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> </p>\n<p>Yes, I’m using GroupKFold for the cross-validation strategy. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1211960,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-02-20T17:46:16.840000",
          "content": "<p><a href=\"https://www.kaggle.com/gianlucarossi\" target=\"_blank\">@gianlucarossi</a>  what are the annotations for ?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1198336,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2021-02-12T23:49:32.297000",
      "content": "<p>0.95 is a very good CV! Did you use GroupKFold for your splits?</p>\n<p>I've just started working on this, and I'm running some tests now to see if the annotations make a difference. Will report back when I have some answers</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1199212,
          "author_name": "datasaurus",
          "author_url": "",
          "post_date": "2021-02-13T16:04:39.227000",
          "content": "<p>I ran 2 experiments using EfficientNet-B0 and 512x512 images, using only <code>train_annotations.csv</code> &amp; 5 folds of GroupKFold.</p>\n<p>Plain Classifer: CV = 0.859<br>\nSegmentation model with classifier head: CV = 0.881</p>\n<p>So for a weak model, adding the annotations certainly helps. I suspect that the gap will close with the larger models, but the question is how much…</p>\n<p>Edit: found an issue with my pipeline - those scores are actually for 256x256</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1199984,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-14T10:05:05.260000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1197237": "Hi,\n\nI'm just getting started in this competition. So far, I've been able to get a 0.95 OOF CV score (GroupKFold) by just using 512x512 images and a relatively lightweight model (EfficientNet B4). I haven't had time to experiment with annotations yet.\n\nI've seen quite a few higher scores on the leaderboard. I'm wondering how important have been annotations for you? How much have they improved your score?\n\nMany thanks!",
    "1201078": "At least in my experience so far I have not seen any value coming from the annotations. Simply seems too sparse of labels and not enough relationship to be found with the things that actually matter, like the heart. Our models actually do quite well given how precise of a detection it is. ",
    "1199954": "Make sure but to ignore `patient_id` in constructing your CV scheme, otherwise your CV results are likely to be badly misleading.",
    "1198519": "Thank you for looking into this @anjum48 \n\nYes, I’m using GroupKFold for the cross-validation strategy. ",
    "1198336": "0.95 is a very good CV! Did you use GroupKFold for your splits?\n\nI've just started working on this, and I'm running some tests now to see if the annotations make a difference. Will report back when I have some answers"
  }
}