{
  "id": 206294,
  "title": "Odd Finding: Weak Neg. ~(-0.31) Correlation btw Data Points to Placement Count",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/206294",
  "author_name": "",
  "post_date": "2020-12-24T01:09:22.401698100Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I was hoping someone could help me understanding this finding. There seems to be a weak negative correlation between the number of data points in an image, <code>len(train_annotations[\"data\"])</code>, and the  number of types of placements in that given image, <code>train.sum(axis=1)</code>.</p>\n<p>This mainly seems odd to me as I would imagine that the more types of irregularities in an image the more data points marking them?</p>\n<p>Any thoughts?</p>",
  "messages": [
    {
      "id": "1124468",
      "postDate": "12/24/2020 01:09:22",
      "content": "<p>I was hoping someone could help me understanding this finding. There seems to be a weak negative correlation between the number of data points in an image, <code>len(train_annotations[\"data\"])</code>, and the  number of types of placements in that given image, <code>train.sum(axis=1)</code>.</p>\n<p>This mainly seems odd to me as I would imagine that the more types of irregularities in an image the more data points marking them?</p>\n<p>Any thoughts?</p>",
      "rawMarkdown": "I was hoping someone could help me understanding this finding. There seems to be a weak negative correlation between the number of data points in an image, `len(train_annotations[\"data\"])`, and the  number of types of placements in that given image, `train.sum(axis=1)`.\n\nThis mainly seems odd to me as I would imagine that the more types of irregularities in an image the more data points marking them?\n\nAny thoughts?",
      "votes": null
    },
    {
      "id": "1124471",
      "postDate": "12/24/2020 01:11:35",
      "content": "<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a> Potentially nothing, but would love your thoughts. Thanks!</p>",
      "rawMarkdown": "jarrelscy Potentially nothing, but would love your thoughts. Thanks!",
      "votes": null
    },
    {
      "id": "1124481",
      "postDate": "12/24/2020 01:32:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dannellyz\" target=\"_blank\">@dannellyz</a> that's cool! I think I've reproduced your observation in (this notebook)[<a href=\"https://www.kaggle.com/jarrelscy/clip-segmentation-visualization\" target=\"_blank\">https://www.kaggle.com/jarrelscy/clip-segmentation-visualization</a>] here. If you have done it the same I did - plot the number of points per line/tube annotation against the total number of line/tubes in the image you can see the negative correlation.</p>\n<p>My guess would be that because labelers label the whole image at one time, when there are many lines and tubes they may be trying to finish it faster and hence putting less points. </p>",
      "rawMarkdown": "Hi @dannellyz that's cool! I think I've reproduced your observation in (this notebook)[https://www.kaggle.com/jarrelscy/clip-segmentation-visualization] here. If you have done it the same I did - plot the number of points per line/tube annotation against the total number of line/tubes in the image you can see the negative correlation.\n\nMy guess would be that because labelers label the whole image at one time, when there are many lines and tubes they may be trying to finish it faster and hence putting less points.",
      "votes": null
    },
    {
      "id": "1124513",
      "postDate": "12/24/2020 02:28:34",
      "content": "<p>Yup thought it maybe something like that. Thanks for the note!</p>",
      "rawMarkdown": "Yup thought it maybe something like that. Thanks for the note!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1124471,
      "author_name": "dannellyz",
      "author_url": "",
      "post_date": "12/24/2020 01:11:35",
      "content": "<p><a href=\"https://www.kaggle.com/jarrelscy\" target=\"_blank\">@jarrelscy</a> Potentially nothing, but would love your thoughts. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1124481,
      "author_name": "jarrelscy",
      "author_url": "",
      "post_date": "12/24/2020 01:32:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/dannellyz\" target=\"_blank\">@dannellyz</a> that's cool! I think I've reproduced your observation in (this notebook)[<a href=\"https://www.kaggle.com/jarrelscy/clip-segmentation-visualization\" target=\"_blank\">https://www.kaggle.com/jarrelscy/clip-segmentation-visualization</a>] here. If you have done it the same I did - plot the number of points per line/tube annotation against the total number of line/tubes in the image you can see the negative correlation.</p>\n<p>My guess would be that because labelers label the whole image at one time, when there are many lines and tubes they may be trying to finish it faster and hence putting less points. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1124513,
          "author_name": "dannellyz",
          "author_url": "",
          "post_date": "12/24/2020 02:28:34",
          "content": "<p>Yup thought it maybe something like that. Thanks for the note!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1124468": "I was hoping someone could help me understanding this finding. There seems to be a weak negative correlation between the number of data points in an image, `len(train_annotations[\"data\"])`, and the  number of types of placements in that given image, `train.sum(axis=1)`.\n\nThis mainly seems odd to me as I would imagine that the more types of irregularities in an image the more data points marking them?\n\nAny thoughts?",
    "1124471": "jarrelscy Potentially nothing, but would love your thoughts. Thanks!",
    "1124481": "Hi @dannellyz that's cool! I think I've reproduced your observation in (this notebook)[https://www.kaggle.com/jarrelscy/clip-segmentation-visualization] here. If you have done it the same I did - plot the number of points per line/tube annotation against the total number of line/tubes in the image you can see the negative correlation.\n\nMy guess would be that because labelers label the whole image at one time, when there are many lines and tubes they may be trying to finish it faster and hence putting less points.",
    "1124513": "Yup thought it maybe something like that. Thanks for the note!"
  },
  "source": "meta"
}