{
  "id": 556420,
  "title": "False positive thyroglobulin detection",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/556420",
  "author_name": "",
  "post_date": "2025-01-13T08:12:36.193183300Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am new to kaggle and am learning a lot from your excellent notebooks and discussions. Thank you.<br>\nI've been looking at the false positives in the model, especially for thyroglobulin, and there are a lot of false positives that are difficult to determine visually. Does this mean that there are some particles similar to thyroglobulin? How should we deal with this kind of problem?</p>\n<h1>&lt;&lt; example of FP &gt;&gt;</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2Fc0996c7dc8a8ce4290b2ad032afe2c58%2F2025-01-13%2017.03.09.png?generation=1736755535825684&amp;alt=media\" alt=\"\"></p>\n<h1>&lt;&lt; example of TP &gt;&gt;</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2F821c13e7572b10f0f75f994c8829c7f7%2F2025-01-13%2017.03.36.png?generation=1736755552696369&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3095323",
      "postDate": "01/13/2025 08:12:36",
      "content": "<p>I am new to kaggle and am learning a lot from your excellent notebooks and discussions. Thank you.<br>\nI've been looking at the false positives in the model, especially for thyroglobulin, and there are a lot of false positives that are difficult to determine visually. Does this mean that there are some particles similar to thyroglobulin? How should we deal with this kind of problem?</p>\n<h1>&lt;&lt; example of FP &gt;&gt;</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2Fc0996c7dc8a8ce4290b2ad032afe2c58%2F2025-01-13%2017.03.09.png?generation=1736755535825684&amp;alt=media\" alt=\"\"></p>\n<h1>&lt;&lt; example of TP &gt;&gt;</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2F821c13e7572b10f0f75f994c8829c7f7%2F2025-01-13%2017.03.36.png?generation=1736755552696369&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I am new to kaggle and am learning a lot from your excellent notebooks and discussions. Thank you.\nI've been looking at the false positives in the model, especially for thyroglobulin, and there are a lot of false positives that are difficult to determine visually. Does this mean that there are some particles similar to thyroglobulin? How should we deal with this kind of problem?\n\n# << example of FP >>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2Fc0996c7dc8a8ce4290b2ad032afe2c58%2F2025-01-13%2017.03.09.png?generation=1736755535825684&alt=media)\n# << example of TP >>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2F821c13e7572b10f0f75f994c8829c7f7%2F2025-01-13%2017.03.36.png?generation=1736755552696369&alt=media)",
      "votes": null
    },
    {
      "id": "3095393",
      "postDate": "01/13/2025 09:50:28",
      "content": "<p>I've seen something similar with my model.  I suspect at least some of them are unlabelled thyroglobulin.  After all, if the competition hosts had a 100% accurate labeling process, why would they need this competition? 😀</p>",
      "rawMarkdown": "I've seen something similar with my model.  I suspect at least some of them are unlabelled thyroglobulin.  After all, if the competition hosts had a 100% accurate labeling process, why would they need this competition? 😀",
      "votes": null
    },
    {
      "id": "3095437",
      "postDate": "01/13/2025 10:58:15",
      "content": "<p>Thanks for your idea. It is difficult when labeling is not always accurate.</p>",
      "rawMarkdown": "Thanks for your idea. It is difficult when labeling is not always accurate.",
      "votes": null
    },
    {
      "id": "3095617",
      "postDate": "01/13/2025 13:53:47",
      "content": "<p>I guess this is why recall and precision have different weights :) </p>",
      "rawMarkdown": "I guess this is why recall and precision have different weights :)",
      "votes": null
    },
    {
      "id": "3098533",
      "postDate": "01/16/2025 15:37:54",
      "content": "<p>I've also seen this with my models. Across all particles, except perhaps the viral particles, I've seen examples of \"false positives\" that look very similar (to the untrained eye at least) to the particle of interest. I've tried (and I'm still trying) hard negative mining to fix this, but so far I have only deteriorated my FBeta score with models trained or retrained with hard negatives specifically included. I personally believe that handling these false positives appropriately will be the key to placing highly in this competition. </p>",
      "rawMarkdown": "I've also seen this with my models. Across all particles, except perhaps the viral particles, I've seen examples of \"false positives\" that look very similar (to the untrained eye at least) to the particle of interest. I've tried (and I'm still trying) hard negative mining to fix this, but so far I have only deteriorated my FBeta score with models trained or retrained with hard negatives specifically included. I personally believe that handling these false positives appropriately will be the key to placing highly in this competition.",
      "votes": null
    },
    {
      "id": "3100250",
      "postDate": "01/19/2025 03:19:13",
      "content": "<p>For what it's worth, this discussion <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/546417\" target=\"_blank\">Ribosome assembly discussion</a> seems to imply that indeed there are true positives that aren't labeled in the data, so it's entirely possible there are \"false positives\" that aren't actually false positives. As <a href=\"https://www.kaggle.com/kvigly55\" target=\"_blank\">@kvigly55</a> pointed out in this thread and as discussed in the linked thread, this is why the competition metric has 16 false positives = 1 false negative.</p>",
      "rawMarkdown": "For what it's worth, this discussion [Ribosome assembly discussion](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/546417) seems to imply that indeed there are true positives that aren't labeled in the data, so it's entirely possible there are \"false positives\" that aren't actually false positives. As @kvigly55 pointed out in this thread and as discussed in the linked thread, this is why the competition metric has 16 false positives = 1 false negative.",
      "votes": null
    },
    {
      "id": "3105873",
      "postDate": "01/24/2025 04:39:06",
      "content": "<p>good idea!</p>",
      "rawMarkdown": "good idea!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3095393,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "01/13/2025 09:50:28",
      "content": "<p>I've seen something similar with my model.  I suspect at least some of them are unlabelled thyroglobulin.  After all, if the competition hosts had a 100% accurate labeling process, why would they need this competition? 😀</p>",
      "votes": null,
      "replies": [
        {
          "id": 3095437,
          "author_name": "yoshinarikawashima",
          "author_url": "",
          "post_date": "01/13/2025 10:58:15",
          "content": "<p>Thanks for your idea. It is difficult when labeling is not always accurate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3095617,
      "author_name": "kvigly55",
      "author_url": "",
      "post_date": "01/13/2025 13:53:47",
      "content": "<p>I guess this is why recall and precision have different weights :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3098533,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "01/16/2025 15:37:54",
      "content": "<p>I've also seen this with my models. Across all particles, except perhaps the viral particles, I've seen examples of \"false positives\" that look very similar (to the untrained eye at least) to the particle of interest. I've tried (and I'm still trying) hard negative mining to fix this, but so far I have only deteriorated my FBeta score with models trained or retrained with hard negatives specifically included. I personally believe that handling these false positives appropriately will be the key to placing highly in this competition. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3100250,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "01/19/2025 03:19:13",
          "content": "<p>For what it's worth, this discussion <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/546417\" target=\"_blank\">Ribosome assembly discussion</a> seems to imply that indeed there are true positives that aren't labeled in the data, so it's entirely possible there are \"false positives\" that aren't actually false positives. As <a href=\"https://www.kaggle.com/kvigly55\" target=\"_blank\">@kvigly55</a> pointed out in this thread and as discussed in the linked thread, this is why the competition metric has 16 false positives = 1 false negative.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3105873,
      "author_name": "kotakatayama",
      "author_url": "",
      "post_date": "01/24/2025 04:39:06",
      "content": "<p>good idea!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3095323": "I am new to kaggle and am learning a lot from your excellent notebooks and discussions. Thank you.\nI've been looking at the false positives in the model, especially for thyroglobulin, and there are a lot of false positives that are difficult to determine visually. Does this mean that there are some particles similar to thyroglobulin? How should we deal with this kind of problem?\n\n# << example of FP >>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2Fc0996c7dc8a8ce4290b2ad032afe2c58%2F2025-01-13%2017.03.09.png?generation=1736755535825684&alt=media)\n# << example of TP >>\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20464683%2F821c13e7572b10f0f75f994c8829c7f7%2F2025-01-13%2017.03.36.png?generation=1736755552696369&alt=media)",
    "3095393": "I've seen something similar with my model.  I suspect at least some of them are unlabelled thyroglobulin.  After all, if the competition hosts had a 100% accurate labeling process, why would they need this competition? 😀",
    "3095437": "Thanks for your idea. It is difficult when labeling is not always accurate.",
    "3095617": "I guess this is why recall and precision have different weights :)",
    "3098533": "I've also seen this with my models. Across all particles, except perhaps the viral particles, I've seen examples of \"false positives\" that look very similar (to the untrained eye at least) to the particle of interest. I've tried (and I'm still trying) hard negative mining to fix this, but so far I have only deteriorated my FBeta score with models trained or retrained with hard negatives specifically included. I personally believe that handling these false positives appropriately will be the key to placing highly in this competition.",
    "3100250": "For what it's worth, this discussion [Ribosome assembly discussion](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/546417) seems to imply that indeed there are true positives that aren't labeled in the data, so it's entirely possible there are \"false positives\" that aren't actually false positives. As @kvigly55 pointed out in this thread and as discussed in the linked thread, this is why the competition metric has 16 false positives = 1 false negative.",
    "3105873": "good idea!"
  },
  "source": "meta"
}