{
  "id": 472212,
  "title": "My optimal CV threshold is close to 0.5, yet my public LB optimal threshold is 0.05",
  "url": "/competitions/blood-vessel-segmentation/discussion/472212",
  "author_name": "",
  "post_date": "2024-01-31T07:32:53.124773100Z",
  "votes": 10,
  "comment_count": 13,
  "views": 0,
  "content": "<p>What does this tell me? What scenario would result in such wildly different optimal thresholds? <br>\nThanks!</p>",
  "messages": [
    {
      "id": "2628246",
      "postDate": "01/31/2024 07:32:53",
      "content": "<p>What does this tell me? What scenario would result in such wildly different optimal thresholds? <br>\nThanks!</p>",
      "rawMarkdown": "What does this tell me? What scenario would result in such wildly different optimal thresholds? \nThanks!",
      "votes": null
    },
    {
      "id": "2628309",
      "postDate": "01/31/2024 08:09:01",
      "content": "<p>single model or ensembles?  single model maybe more sensitive .</p>",
      "rawMarkdown": "single model or ensembles?  single model maybe more sensitive .",
      "votes": null
    },
    {
      "id": "2628336",
      "postDate": "01/31/2024 08:35:58",
      "content": "<p>Single model. Sensitive to what? I mean what's the mechanism behind it? Why the model underpredicts so much? I have never experienced such a massive AND consistent difference between LB and CV thresholds before.</p>",
      "rawMarkdown": "Single model. Sensitive to what? I mean what's the mechanism behind it? Why the model underpredicts so much? I have never experienced such a massive AND consistent difference between LB and CV thresholds before.",
      "votes": null
    },
    {
      "id": "2628377",
      "postDate": "01/31/2024 09:04:22",
      "content": "<p>It means you overfitted the public leaderboard! 😅</p>",
      "rawMarkdown": "It means you overfitted the public leaderboard! 😅",
      "votes": null
    },
    {
      "id": "2628401",
      "postDate": "01/31/2024 09:27:15",
      "content": "<p>I know, but I want to understand what causes such huge threshold discrepancies between two seemingly similar images. Like I usually see something like 0.35-0.65 optimal threshold range. But here it's a 10x difference.</p>",
      "rawMarkdown": "I know, but I want to understand what causes such huge threshold discrepancies between two seemingly similar images. Like I usually see something like 0.35-0.65 optimal threshold range. But here it's a 10x difference.",
      "votes": null
    },
    {
      "id": "2628417",
      "postDate": "01/31/2024 09:37:58",
      "content": "<p>I see a few possibilities:</p>\n<ul>\n<li><strong>different annotation software/hyper parameters for ground truth</strong>: the competition metric is very sensitive (at the pixel level).</li>\n<li><strong>difference in labels sparsity</strong>: even the dense kidneys seems to be missing labels IMO, if for some reason the public kidney is perfectly segmented, then lowering your threshold will probably improve the score as your model will be rewarded when detecting small vessels.</li>\n<li><strong>data shift</strong>: the technology seems quite new and all stacks seem to have different type of noise, SNR etc… Again, the competition metric is very sensitive.</li>\n</ul>",
      "rawMarkdown": "I see a few possibilities:\n- **different annotation software/hyper parameters for ground truth**: the competition metric is very sensitive (at the pixel level).\n- **difference in labels sparsity**: even the dense kidneys seems to be missing labels IMO, if for some reason the public kidney is perfectly segmented, then lowering your threshold will probably improve the score as your model will be rewarded when detecting small vessels.\n- **data shift**: the technology seems quite new and all stacks seem to have different type of noise, SNR etc... Again, the competition metric is very sensitive.",
      "votes": null
    },
    {
      "id": "2628638",
      "postDate": "01/31/2024 12:25:01",
      "content": "<p>BUG?</p>\n<p>can't tell unless we know the score values. if the score values differs by too much, could be bug.<br>\nif there is no bug, here is the reason:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F33bc2aabb91c7ed3c0a4922f63758776%2FSelection_999(4722).png?generation=1706703897494348&amp;alt=media\"></p>",
      "rawMarkdown": "BUG?\n\ncan't tell unless we know the score values. if the score values differs by too much, could be bug.\nif there is no bug, here is the reason:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F33bc2aabb91c7ed3c0a4922f63758776%2FSelection_999(4722).png?generation=1706703897494348&alt=media)",
      "votes": null
    },
    {
      "id": "2628704",
      "postDate": "01/31/2024 12:55:47",
      "content": "<p>Not much of a difference in absolute values, only in thresholds<br>\nCV @ th 0.5 = 0.875<br>\nLB @ th 0.05 = 0.870</p>",
      "rawMarkdown": "Not much of a difference in absolute values, only in thresholds\nCV @ th 0.5 = 0.875\nLB @ th 0.05 = 0.870",
      "votes": null
    },
    {
      "id": "2628705",
      "postDate": "01/31/2024 12:57:13",
      "content": "<p>Is that on kidney 3 dense ?<br>\nAlso what is your LB score <a href=\"https://www.kaggle.com/0.5\" target=\"_blank\">@0.5</a>?</p>",
      "rawMarkdown": "Is that on kidney 3 dense ?\nAlso what is your LB score @0.5?",
      "votes": null
    },
    {
      "id": "2629162",
      "postDate": "01/31/2024 16:56:00",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb55c3f75d514baca24e9a8008a74c258%2FSelection_999(4731).png?generation=1706720105023319&amp;alt=media\"></p>\n<p>we are more used to region dice.<br>\nfor boundary dice, it is different.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb55c3f75d514baca24e9a8008a74c258%2FSelection_999(4731).png?generation=1706720105023319&alt=media)\n\nwe are more used to region dice.\nfor boundary dice, it is different.",
      "votes": null
    },
    {
      "id": "2629723",
      "postDate": "02/01/2024 01:45:38",
      "content": "<p>If it's 0.5, what's the public score? These are some local results I have, which show different patterns. I believe, besides the data itself, it's greatly related to our model architecture and data augmentation.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F16a9585642af5d3a685307435b8350bf%2F5741706751861_.pic.jpg?generation=1706751937233619&amp;alt=media\"></p>",
      "rawMarkdown": "If it's 0.5, what's the public score? These are some local results I have, which show different patterns. I believe, besides the data itself, it's greatly related to our model architecture and data augmentation.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F16a9585642af5d3a685307435b8350bf%2F5741706751861_.pic.jpg?generation=1706751937233619&alt=media)",
      "votes": null
    },
    {
      "id": "2630038",
      "postDate": "02/01/2024 06:02:28",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> <br>\nFor me, such differences in the threshold value for the public lb gives really different scores. Do you have the same case?</p>",
      "rawMarkdown": "lihaoweicvch \nFor me, such differences in the threshold value for the public lb gives really different scores. Do you have the same case?",
      "votes": null
    },
    {
      "id": "2630055",
      "postDate": "02/01/2024 06:14:08",
      "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>  My single convolution based model is the same case, ensemble or ViT are much better </p>",
      "rawMarkdown": "mohammad2012191  My single convolution based model is the same case, ensemble or ViT are much better",
      "votes": null
    },
    {
      "id": "2631958",
      "postDate": "02/02/2024 04:18:42",
      "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> </p>\n<p>These different patterns look interesting! What differences in the architecture, augmentations or other may have caused these?</p>",
      "rawMarkdown": "lihaoweicvch \n\nThese different patterns look interesting! What differences in the architecture, augmentations or other may have caused these?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2628309,
      "author_name": "lihaoweicvch",
      "author_url": "",
      "post_date": "01/31/2024 08:09:01",
      "content": "<p>single model or ensembles?  single model maybe more sensitive .</p>",
      "votes": null,
      "replies": [
        {
          "id": 2628336,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "01/31/2024 08:35:58",
          "content": "<p>Single model. Sensitive to what? I mean what's the mechanism behind it? Why the model underpredicts so much? I have never experienced such a massive AND consistent difference between LB and CV thresholds before.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2628638,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "01/31/2024 12:25:01",
              "content": "<p>BUG?</p>\n<p>can't tell unless we know the score values. if the score values differs by too much, could be bug.<br>\nif there is no bug, here is the reason:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F33bc2aabb91c7ed3c0a4922f63758776%2FSelection_999(4722).png?generation=1706703897494348&amp;alt=media\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2628704,
                  "author_name": "sakvaua",
                  "author_url": "",
                  "post_date": "01/31/2024 12:55:47",
                  "content": "<p>Not much of a difference in absolute values, only in thresholds<br>\nCV @ th 0.5 = 0.875<br>\nLB @ th 0.05 = 0.870</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2628705,
                      "author_name": "optimo",
                      "author_url": "",
                      "post_date": "01/31/2024 12:57:13",
                      "content": "<p>Is that on kidney 3 dense ?<br>\nAlso what is your LB score <a href=\"https://www.kaggle.com/0.5\" target=\"_blank\">@0.5</a>?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2629162,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "01/31/2024 16:56:00",
                          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb55c3f75d514baca24e9a8008a74c258%2FSelection_999(4731).png?generation=1706720105023319&amp;alt=media\"></p>\n<p>we are more used to region dice.<br>\nfor boundary dice, it is different.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 2629723,
              "author_name": "lihaoweicvch",
              "author_url": "",
              "post_date": "02/01/2024 01:45:38",
              "content": "<p>If it's 0.5, what's the public score? These are some local results I have, which show different patterns. I believe, besides the data itself, it's greatly related to our model architecture and data augmentation.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F16a9585642af5d3a685307435b8350bf%2F5741706751861_.pic.jpg?generation=1706751937233619&amp;alt=media\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2630038,
                  "author_name": "mohammad2012191",
                  "author_url": "",
                  "post_date": "02/01/2024 06:02:28",
                  "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> <br>\nFor me, such differences in the threshold value for the public lb gives really different scores. Do you have the same case?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2630055,
                      "author_name": "lihaoweicvch",
                      "author_url": "",
                      "post_date": "02/01/2024 06:14:08",
                      "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a>  My single convolution based model is the same case, ensemble or ViT are much better </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                },
                {
                  "id": 2631958,
                  "author_name": "crustacean",
                  "author_url": "",
                  "post_date": "02/02/2024 04:18:42",
                  "content": "<p><a href=\"https://www.kaggle.com/lihaoweicvch\" target=\"_blank\">@lihaoweicvch</a> </p>\n<p>These different patterns look interesting! What differences in the architecture, augmentations or other may have caused these?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2628377,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "01/31/2024 09:04:22",
      "content": "<p>It means you overfitted the public leaderboard! 😅</p>",
      "votes": null,
      "replies": [
        {
          "id": 2628401,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "01/31/2024 09:27:15",
          "content": "<p>I know, but I want to understand what causes such huge threshold discrepancies between two seemingly similar images. Like I usually see something like 0.35-0.65 optimal threshold range. But here it's a 10x difference.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2628417,
              "author_name": "optimo",
              "author_url": "",
              "post_date": "01/31/2024 09:37:58",
              "content": "<p>I see a few possibilities:</p>\n<ul>\n<li><strong>different annotation software/hyper parameters for ground truth</strong>: the competition metric is very sensitive (at the pixel level).</li>\n<li><strong>difference in labels sparsity</strong>: even the dense kidneys seems to be missing labels IMO, if for some reason the public kidney is perfectly segmented, then lowering your threshold will probably improve the score as your model will be rewarded when detecting small vessels.</li>\n<li><strong>data shift</strong>: the technology seems quite new and all stacks seem to have different type of noise, SNR etc… Again, the competition metric is very sensitive.</li>\n</ul>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2628246": "What does this tell me? What scenario would result in such wildly different optimal thresholds? \nThanks!",
    "2628309": "single model or ensembles?  single model maybe more sensitive .",
    "2628336": "Single model. Sensitive to what? I mean what's the mechanism behind it? Why the model underpredicts so much? I have never experienced such a massive AND consistent difference between LB and CV thresholds before.",
    "2628377": "It means you overfitted the public leaderboard! 😅",
    "2628401": "I know, but I want to understand what causes such huge threshold discrepancies between two seemingly similar images. Like I usually see something like 0.35-0.65 optimal threshold range. But here it's a 10x difference.",
    "2628417": "I see a few possibilities:\n- **different annotation software/hyper parameters for ground truth**: the competition metric is very sensitive (at the pixel level).\n- **difference in labels sparsity**: even the dense kidneys seems to be missing labels IMO, if for some reason the public kidney is perfectly segmented, then lowering your threshold will probably improve the score as your model will be rewarded when detecting small vessels.\n- **data shift**: the technology seems quite new and all stacks seem to have different type of noise, SNR etc... Again, the competition metric is very sensitive.",
    "2628638": "BUG?\n\ncan't tell unless we know the score values. if the score values differs by too much, could be bug.\nif there is no bug, here is the reason:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F33bc2aabb91c7ed3c0a4922f63758776%2FSelection_999(4722).png?generation=1706703897494348&alt=media)",
    "2628704": "Not much of a difference in absolute values, only in thresholds\nCV @ th 0.5 = 0.875\nLB @ th 0.05 = 0.870",
    "2628705": "Is that on kidney 3 dense ?\nAlso what is your LB score @0.5?",
    "2629162": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb55c3f75d514baca24e9a8008a74c258%2FSelection_999(4731).png?generation=1706720105023319&alt=media)\n\nwe are more used to region dice.\nfor boundary dice, it is different.",
    "2629723": "If it's 0.5, what's the public score? These are some local results I have, which show different patterns. I believe, besides the data itself, it's greatly related to our model architecture and data augmentation.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F16a9585642af5d3a685307435b8350bf%2F5741706751861_.pic.jpg?generation=1706751937233619&alt=media)",
    "2630038": "lihaoweicvch \nFor me, such differences in the threshold value for the public lb gives really different scores. Do you have the same case?",
    "2630055": "mohammad2012191  My single convolution based model is the same case, ensemble or ViT are much better",
    "2631958": "lihaoweicvch \n\nThese different patterns look interesting! What differences in the architecture, augmentations or other may have caused these?"
  },
  "source": "meta"
}