{
  "id": 216219,
  "title": "Why CVC is actually your best performing classifier",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/216219",
  "author_name": "",
  "post_date": "2021-02-02T04:00:27.864124600Z",
  "votes": 60,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I spent a bit of time looking at the errors the model is making. Slicing up the error patterns a few different ways, looking at error of each group individually I see that performance looks like this: </p>\n<table>\n<thead>\n<tr>\n<th>ETT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ETT - Abnormal</td>\n<td>0.962</td>\n</tr>\n<tr>\n<td>ETT - Borderline</td>\n<td>0.9538</td>\n</tr>\n<tr>\n<td>ETT - Normal</td>\n<td>0.9901</td>\n</tr>\n<tr>\n<td><strong>ETT AVG</strong></td>\n<td><strong>0.9686333333333333</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>NGT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NGT - Abnormal</td>\n<td>0.9381</td>\n</tr>\n<tr>\n<td>NGT - Borderline</td>\n<td>0.9491</td>\n</tr>\n<tr>\n<td>NGT - Incompletely Imaged</td>\n<td>0.9797</td>\n</tr>\n<tr>\n<td>NGT - Normal</td>\n<td>0.9837</td>\n</tr>\n<tr>\n<td><strong>NGT AVG</strong></td>\n<td><strong>0.96265</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>CVC Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CVC - Abnormal</td>\n<td>0.9138</td>\n</tr>\n<tr>\n<td>CVC - Borderline</td>\n<td>0.8464</td>\n</tr>\n<tr>\n<td>CVC - Normal</td>\n<td>0.9081</td>\n</tr>\n<tr>\n<td><strong>CVC AVG</strong></td>\n<td><strong>0.8894333333333333</strong></td>\n</tr>\n</tbody>\n</table>\n<p>This tells us the story that CVC is the hardest to classify with significantly worse average AUC than the other types of catheters. Looking a bit more closely we will find that that is not true. </p>\n<p>If we look only at the samples that have the actual catheter group in question then we find significantly different performance. That's to say, of the ~30k X-rays only 8457 have ETT and 8308 NGT while nearly every single X-ray has the CVC type. This means that the AUC of NGT and ETT are boosted simply by outputting zeros when those types are not present which is a much easier thing to determine. Looking at these subsets we see</p>\n<table>\n<thead>\n<tr>\n<th>ETT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ETT - Abnormal</td>\n<td>0.935</td>\n</tr>\n<tr>\n<td>ETT - Borderline</td>\n<td>0.8436</td>\n</tr>\n<tr>\n<td>ETT - Normal</td>\n<td>0.8567</td>\n</tr>\n<tr>\n<td><strong>ETT AVG</strong></td>\n<td><strong>0.8784333333333333</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>NGT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NGT - Abnormal</td>\n<td>0.8614</td>\n</tr>\n<tr>\n<td>NGT - Borderline</td>\n<td>0.8508</td>\n</tr>\n<tr>\n<td>NGT - Incompletely Imaged</td>\n<td>0.9175</td>\n</tr>\n<tr>\n<td>NGT - Normal</td>\n<td>0.9024</td>\n</tr>\n<tr>\n<td><strong>NGT AVG</strong></td>\n<td><strong>0.8830250000000001</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>CVC Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CVC - Abnormal</td>\n<td>0.915</td>\n</tr>\n<tr>\n<td>CVC - Borderline</td>\n<td>0.8473</td>\n</tr>\n<tr>\n<td>CVC - Normal</td>\n<td>0.9047</td>\n</tr>\n<tr>\n<td><strong>CVC AVG</strong></td>\n<td><strong>0.8890000000000001</strong></td>\n</tr>\n</tbody>\n</table>\n<p>All of a sudden we see that ETT and NGT when evaluated on X-rays that actually have those types the performance is even below that of CVC. The free points from simply detecting if the catheter is present are excluded from this view of the error. Since almost all X-rays have the CVC line the performance does not move very much. </p>",
  "messages": [
    {
      "id": "1181678",
      "postDate": "02/02/2021 04:00:27",
      "content": "<p>I spent a bit of time looking at the errors the model is making. Slicing up the error patterns a few different ways, looking at error of each group individually I see that performance looks like this: </p>\n<table>\n<thead>\n<tr>\n<th>ETT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ETT - Abnormal</td>\n<td>0.962</td>\n</tr>\n<tr>\n<td>ETT - Borderline</td>\n<td>0.9538</td>\n</tr>\n<tr>\n<td>ETT - Normal</td>\n<td>0.9901</td>\n</tr>\n<tr>\n<td><strong>ETT AVG</strong></td>\n<td><strong>0.9686333333333333</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>NGT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NGT - Abnormal</td>\n<td>0.9381</td>\n</tr>\n<tr>\n<td>NGT - Borderline</td>\n<td>0.9491</td>\n</tr>\n<tr>\n<td>NGT - Incompletely Imaged</td>\n<td>0.9797</td>\n</tr>\n<tr>\n<td>NGT - Normal</td>\n<td>0.9837</td>\n</tr>\n<tr>\n<td><strong>NGT AVG</strong></td>\n<td><strong>0.96265</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>CVC Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CVC - Abnormal</td>\n<td>0.9138</td>\n</tr>\n<tr>\n<td>CVC - Borderline</td>\n<td>0.8464</td>\n</tr>\n<tr>\n<td>CVC - Normal</td>\n<td>0.9081</td>\n</tr>\n<tr>\n<td><strong>CVC AVG</strong></td>\n<td><strong>0.8894333333333333</strong></td>\n</tr>\n</tbody>\n</table>\n<p>This tells us the story that CVC is the hardest to classify with significantly worse average AUC than the other types of catheters. Looking a bit more closely we will find that that is not true. </p>\n<p>If we look only at the samples that have the actual catheter group in question then we find significantly different performance. That's to say, of the ~30k X-rays only 8457 have ETT and 8308 NGT while nearly every single X-ray has the CVC type. This means that the AUC of NGT and ETT are boosted simply by outputting zeros when those types are not present which is a much easier thing to determine. Looking at these subsets we see</p>\n<table>\n<thead>\n<tr>\n<th>ETT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ETT - Abnormal</td>\n<td>0.935</td>\n</tr>\n<tr>\n<td>ETT - Borderline</td>\n<td>0.8436</td>\n</tr>\n<tr>\n<td>ETT - Normal</td>\n<td>0.8567</td>\n</tr>\n<tr>\n<td><strong>ETT AVG</strong></td>\n<td><strong>0.8784333333333333</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>NGT Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NGT - Abnormal</td>\n<td>0.8614</td>\n</tr>\n<tr>\n<td>NGT - Borderline</td>\n<td>0.8508</td>\n</tr>\n<tr>\n<td>NGT - Incompletely Imaged</td>\n<td>0.9175</td>\n</tr>\n<tr>\n<td>NGT - Normal</td>\n<td>0.9024</td>\n</tr>\n<tr>\n<td><strong>NGT AVG</strong></td>\n<td><strong>0.8830250000000001</strong></td>\n</tr>\n</tbody>\n</table>\n<table>\n<thead>\n<tr>\n<th>CVC Error</th>\n<th>AUC</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>CVC - Abnormal</td>\n<td>0.915</td>\n</tr>\n<tr>\n<td>CVC - Borderline</td>\n<td>0.8473</td>\n</tr>\n<tr>\n<td>CVC - Normal</td>\n<td>0.9047</td>\n</tr>\n<tr>\n<td><strong>CVC AVG</strong></td>\n<td><strong>0.8890000000000001</strong></td>\n</tr>\n</tbody>\n</table>\n<p>All of a sudden we see that ETT and NGT when evaluated on X-rays that actually have those types the performance is even below that of CVC. The free points from simply detecting if the catheter is present are excluded from this view of the error. Since almost all X-rays have the CVC line the performance does not move very much. </p>",
      "rawMarkdown": "I spent a bit of time looking at the errors the model is making. Slicing up the error patterns a few different ways, looking at error of each group individually I see that performance looks like this: \n\n| ETT Error        | AUC                |\n|------------------|--------------------|\n| ETT - Abnormal   | 0.962              |\n| ETT - Borderline | 0.9538             |\n| ETT - Normal     | 0.9901             |\n| **ETT AVG**          | **0.9686333333333333** |\n\n\n\n| NGT Error                 | AUC     |\n|---------------------------|---------|\n| NGT - Abnormal            | 0.9381  |\n| NGT - Borderline          | 0.9491  |\n| NGT - Incompletely Imaged | 0.9797  |\n| NGT - Normal              | 0.9837  |\n| **NGT AVG**                  | **0.96265** |\n\n\n| CVC Error        | AUC                |\n|------------------|--------------------|\n| CVC - Abnormal   | 0.9138             |\n| CVC - Borderline | 0.8464             |\n| CVC - Normal     | 0.9081             |\n| **CVC AVG**          | **0.8894333333333333** |\n\nThis tells us the story that CVC is the hardest to classify with significantly worse average AUC than the other types of catheters. Looking a bit more closely we will find that that is not true. \n\nIf we look only at the samples that have the actual catheter group in question then we find significantly different performance. That's to say, of the ~30k X-rays only 8457 have ETT and 8308 NGT while nearly every single X-ray has the CVC type. This means that the AUC of NGT and ETT are boosted simply by outputting zeros when those types are not present which is a much easier thing to determine. Looking at these subsets we see\n\n| ETT Error        | AUC                |\n|------------------|--------------------|\n| ETT - Abnormal   | 0.935              |\n| ETT - Borderline | 0.8436             |\n| ETT - Normal     | 0.8567             |\n| **ETT AVG**          | **0.8784333333333333** |\n\n\n| NGT Error                 | AUC                |\n|---------------------------|--------------------|\n| NGT - Abnormal            | 0.8614             |\n| NGT - Borderline          | 0.8508             |\n| NGT - Incompletely Imaged | 0.9175             |\n| NGT - Normal              | 0.9024             |\n| **NGT AVG**                   | **0.8830250000000001** |\n  \n\n| CVC Error        | AUC                |\n|------------------|--------------------|\n| CVC - Abnormal   | 0.915              |\n| CVC - Borderline | 0.8473             |\n| CVC - Normal     | 0.9047             |\n| **CVC AVG**          | **0.8890000000000001** |\n\nAll of a sudden we see that ETT and NGT when evaluated on X-rays that actually have those types the performance is even below that of CVC. The free points from simply detecting if the catheter is present are excluded from this view of the error. Since almost all X-rays have the CVC line the performance does not move very much.",
      "votes": null
    },
    {
      "id": "1181709",
      "postDate": "02/02/2021 04:34:40",
      "content": "<p>Looking at the histograms of the predictions we can see that the spread of the certainty of the CVC is much higher than that of ETT, but that is mostly because it is trivial for the model to separate between when the catheters are present or not while the CVC classification always needs to categorize into the various groups.</p>\n<p><img src=\"https://i.imgur.com/inrA4r9.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/432eqWX.png\" alt=\"\"></p>",
      "rawMarkdown": "Looking at the histograms of the predictions we can see that the spread of the certainty of the CVC is much higher than that of ETT, but that is mostly because it is trivial for the model to separate between when the catheters are present or not while the CVC classification always needs to categorize into the various groups.\n\n![](https://i.imgur.com/inrA4r9.png)\n\n![](https://i.imgur.com/432eqWX.png)",
      "votes": null
    },
    {
      "id": "1181720",
      "postDate": "02/02/2021 04:40:42",
      "content": "<p><img src=\"https://i.imgur.com/UqLImd4.png\" alt=\"\"></p>\n<p>Looking at the number of samples with more than 1 label within a category we see that there are none in ETT, only 45 in NGT and 3575 in CVC. So in theory we could do this as a single classification with softmax function instead of sigmoid function and multilabel classification but there are some downsides to this.</p>\n<p>Using a softmax instead of sigmoid might yield better classification but what we really care about is the sort order of the predictions, applying the softmax makes it so the confidence of the predictions of the other columns also impacts the prediction within the column so the sort might not be as consistent. </p>",
      "rawMarkdown": "![](https://i.imgur.com/UqLImd4.png)\n\nLooking at the number of samples with more than 1 label within a category we see that there are none in ETT, only 45 in NGT and 3575 in CVC. So in theory we could do this as a single classification with softmax function instead of sigmoid function and multilabel classification but there are some downsides to this.\n\nUsing a softmax instead of sigmoid might yield better classification but what we really care about is the sort order of the predictions, applying the softmax makes it so the confidence of the predictions of the other columns also impacts the prediction within the column so the sort might not be as consistent.",
      "votes": null
    },
    {
      "id": "1181757",
      "postDate": "02/02/2021 05:23:04",
      "content": "<p>In addition for CVC think only around 71 have 3 the rest of the multi CVCs are 2. For NGT the 45 are only 2 rest are 1.  Started looking at OOF analysis and realised this multi issue and considering how to address. Was thinking if separate CVC model might be useful or adding a present or not target for each of these ETT, NGT, CVC like Swan Ganz Catheter Present.  For Swan and CVC there are some issues noted by raddar in the Welcome post still awaiting organiser follow up - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203367#1138150\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203367#1138150</a>  It seems if using annotations only Swan is labelled, without some/all have CVC Normal too.  </p>",
      "rawMarkdown": "In addition for CVC think only around 71 have 3 the rest of the multi CVCs are 2. For NGT the 45 are only 2 rest are 1.  Started looking at OOF analysis and realised this multi issue and considering how to address. Was thinking if separate CVC model might be useful or adding a present or not target for each of these ETT, NGT, CVC like Swan Ganz Catheter Present.  For Swan and CVC there are some issues noted by raddar in the Welcome post still awaiting organiser follow up - https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203367#1138150  It seems if using annotations only Swan is labelled, without some/all have CVC Normal too.",
      "votes": null
    },
    {
      "id": "1181991",
      "postDate": "02/02/2021 08:37:14",
      "content": "<p>Love your post. I hated AUC metric from the start of the competition.. The thing that competition organizers are most interested in is \"Abnormal\" classes anyway and AUC is not an ideal proxy metric for the real problem</p>",
      "rawMarkdown": "Love your post. I hated AUC metric from the start of the competition.. The thing that competition organizers are most interested in is \"Abnormal\" classes anyway and AUC is not an ideal proxy metric for the real problem",
      "votes": null
    },
    {
      "id": "1182092",
      "postDate": "02/02/2021 09:50:24",
      "content": "<p>I think at least auc free us from tuning threshold and developing post-processing methods.<br>\nWhich metric is better for this real problem from your view?</p>",
      "rawMarkdown": "I think at least auc free us from tuning threshold and developing post-processing methods.\nWhich metric is better for this real problem from your view?",
      "votes": null
    },
    {
      "id": "1182143",
      "postDate": "02/02/2021 10:20:47",
      "content": "<p>I nominate F-beta</p>",
      "rawMarkdown": "I nominate F-beta",
      "votes": null
    },
    {
      "id": "1187146",
      "postDate": "02/05/2021 08:35:17",
      "content": "<p>FYI - there are 24 train entries with no targets at all.  </p>",
      "rawMarkdown": "FYI - there are 24 train entries with no targets at all.",
      "votes": null
    },
    {
      "id": "1216966",
      "postDate": "02/24/2021 16:56:04",
      "content": "<p>Nice analysis. Thanks for sharing.</p>",
      "rawMarkdown": "Nice analysis. Thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1181709,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/02/2021 04:34:40",
      "content": "<p>Looking at the histograms of the predictions we can see that the spread of the certainty of the CVC is much higher than that of ETT, but that is mostly because it is trivial for the model to separate between when the catheters are present or not while the CVC classification always needs to categorize into the various groups.</p>\n<p><img src=\"https://i.imgur.com/inrA4r9.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/432eqWX.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1181720,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/02/2021 04:40:42",
          "content": "<p><img src=\"https://i.imgur.com/UqLImd4.png\" alt=\"\"></p>\n<p>Looking at the number of samples with more than 1 label within a category we see that there are none in ETT, only 45 in NGT and 3575 in CVC. So in theory we could do this as a single classification with softmax function instead of sigmoid function and multilabel classification but there are some downsides to this.</p>\n<p>Using a softmax instead of sigmoid might yield better classification but what we really care about is the sort order of the predictions, applying the softmax makes it so the confidence of the predictions of the other columns also impacts the prediction within the column so the sort might not be as consistent. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1181757,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "02/02/2021 05:23:04",
          "content": "<p>In addition for CVC think only around 71 have 3 the rest of the multi CVCs are 2. For NGT the 45 are only 2 rest are 1.  Started looking at OOF analysis and realised this multi issue and considering how to address. Was thinking if separate CVC model might be useful or adding a present or not target for each of these ETT, NGT, CVC like Swan Ganz Catheter Present.  For Swan and CVC there are some issues noted by raddar in the Welcome post still awaiting organiser follow up - <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203367#1138150\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203367#1138150</a>  It seems if using annotations only Swan is labelled, without some/all have CVC Normal too.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1187146,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "02/05/2021 08:35:17",
          "content": "<p>FYI - there are 24 train entries with no targets at all.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1181991,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "02/02/2021 08:37:14",
      "content": "<p>Love your post. I hated AUC metric from the start of the competition.. The thing that competition organizers are most interested in is \"Abnormal\" classes anyway and AUC is not an ideal proxy metric for the real problem</p>",
      "votes": null,
      "replies": [
        {
          "id": 1182092,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/02/2021 09:50:24",
          "content": "<p>I think at least auc free us from tuning threshold and developing post-processing methods.<br>\nWhich metric is better for this real problem from your view?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1182143,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "02/02/2021 10:20:47",
          "content": "<p>I nominate F-beta</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1216966,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/24/2021 16:56:04",
      "content": "<p>Nice analysis. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1181678": "I spent a bit of time looking at the errors the model is making. Slicing up the error patterns a few different ways, looking at error of each group individually I see that performance looks like this: \n\n| ETT Error        | AUC                |\n|------------------|--------------------|\n| ETT - Abnormal   | 0.962              |\n| ETT - Borderline | 0.9538             |\n| ETT - Normal     | 0.9901             |\n| **ETT AVG**          | **0.9686333333333333** |\n\n\n\n| NGT Error                 | AUC     |\n|---------------------------|---------|\n| NGT - Abnormal            | 0.9381  |\n| NGT - Borderline          | 0.9491  |\n| NGT - Incompletely Imaged | 0.9797  |\n| NGT - Normal              | 0.9837  |\n| **NGT AVG**                  | **0.96265** |\n\n\n| CVC Error        | AUC                |\n|------------------|--------------------|\n| CVC - Abnormal   | 0.9138             |\n| CVC - Borderline | 0.8464             |\n| CVC - Normal     | 0.9081             |\n| **CVC AVG**          | **0.8894333333333333** |\n\nThis tells us the story that CVC is the hardest to classify with significantly worse average AUC than the other types of catheters. Looking a bit more closely we will find that that is not true. \n\nIf we look only at the samples that have the actual catheter group in question then we find significantly different performance. That's to say, of the ~30k X-rays only 8457 have ETT and 8308 NGT while nearly every single X-ray has the CVC type. This means that the AUC of NGT and ETT are boosted simply by outputting zeros when those types are not present which is a much easier thing to determine. Looking at these subsets we see\n\n| ETT Error        | AUC                |\n|------------------|--------------------|\n| ETT - Abnormal   | 0.935              |\n| ETT - Borderline | 0.8436             |\n| ETT - Normal     | 0.8567             |\n| **ETT AVG**          | **0.8784333333333333** |\n\n\n| NGT Error                 | AUC                |\n|---------------------------|--------------------|\n| NGT - Abnormal            | 0.8614             |\n| NGT - Borderline          | 0.8508             |\n| NGT - Incompletely Imaged | 0.9175             |\n| NGT - Normal              | 0.9024             |\n| **NGT AVG**                   | **0.8830250000000001** |\n  \n\n| CVC Error        | AUC                |\n|------------------|--------------------|\n| CVC - Abnormal   | 0.915              |\n| CVC - Borderline | 0.8473             |\n| CVC - Normal     | 0.9047             |\n| **CVC AVG**          | **0.8890000000000001** |\n\nAll of a sudden we see that ETT and NGT when evaluated on X-rays that actually have those types the performance is even below that of CVC. The free points from simply detecting if the catheter is present are excluded from this view of the error. Since almost all X-rays have the CVC line the performance does not move very much.",
    "1181709": "Looking at the histograms of the predictions we can see that the spread of the certainty of the CVC is much higher than that of ETT, but that is mostly because it is trivial for the model to separate between when the catheters are present or not while the CVC classification always needs to categorize into the various groups.\n\n![](https://i.imgur.com/inrA4r9.png)\n\n![](https://i.imgur.com/432eqWX.png)",
    "1181720": "![](https://i.imgur.com/UqLImd4.png)\n\nLooking at the number of samples with more than 1 label within a category we see that there are none in ETT, only 45 in NGT and 3575 in CVC. So in theory we could do this as a single classification with softmax function instead of sigmoid function and multilabel classification but there are some downsides to this.\n\nUsing a softmax instead of sigmoid might yield better classification but what we really care about is the sort order of the predictions, applying the softmax makes it so the confidence of the predictions of the other columns also impacts the prediction within the column so the sort might not be as consistent.",
    "1181757": "In addition for CVC think only around 71 have 3 the rest of the multi CVCs are 2. For NGT the 45 are only 2 rest are 1.  Started looking at OOF analysis and realised this multi issue and considering how to address. Was thinking if separate CVC model might be useful or adding a present or not target for each of these ETT, NGT, CVC like Swan Ganz Catheter Present.  For Swan and CVC there are some issues noted by raddar in the Welcome post still awaiting organiser follow up - https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/203367#1138150  It seems if using annotations only Swan is labelled, without some/all have CVC Normal too.",
    "1181991": "Love your post. I hated AUC metric from the start of the competition.. The thing that competition organizers are most interested in is \"Abnormal\" classes anyway and AUC is not an ideal proxy metric for the real problem",
    "1182092": "I think at least auc free us from tuning threshold and developing post-processing methods.\nWhich metric is better for this real problem from your view?",
    "1182143": "I nominate F-beta",
    "1187146": "FYI - there are 24 train entries with no targets at all.",
    "1216966": "Nice analysis. Thanks for sharing."
  },
  "source": "meta"
}