{
  "id": 480452,
  "title": "What is Wrong with Private Data ?",
  "url": "/competitions/blood-vessel-segmentation/discussion/480452",
  "author_name": "",
  "post_date": "2024-02-28T15:52:11.235657400Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm a bit late to the party, but I decided to rerun my model with different thresholds for converting probabilities to masks, to assess how sensitive to it private LB is.<br>\nLocally, my optimal thresholds were between 0.3 and 0.5 for kidneys 1/2/3. Public LB scored quite well with th=0.3 but it's not the case for private … </p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>Private Score</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>th=0.01</td>\n<td>0.733862</td>\n<td>0.819565</td>\n</tr>\n<tr>\n<td>th=0.025</td>\n<td><strong>0.738578</strong></td>\n<td>0.826613</td>\n</tr>\n<tr>\n<td>th=0.05</td>\n<td>0.71296</td>\n<td>0.828603</td>\n</tr>\n<tr>\n<td>th=0.1</td>\n<td>0.678819</td>\n<td>0.829803</td>\n</tr>\n<tr>\n<td>th=0.15</td>\n<td>0.658714</td>\n<td>0.830981</td>\n</tr>\n<tr>\n<td>th=0.2</td>\n<td>0.643396</td>\n<td><strong>0.831588</strong></td>\n</tr>\n<tr>\n<td>th=0.25</td>\n<td>0.629126</td>\n<td>0.829168</td>\n</tr>\n<tr>\n<td>th=0.3</td>\n<td>0.616352</td>\n<td>0.826501</td>\n</tr>\n<tr>\n<td>th=0.5</td>\n<td>0.567844</td>\n<td>0.806398</td>\n</tr>\n</tbody>\n</table>\n<p>The table speaks for itself. This model ranks #3 with an absurdly low threshold of 0.025. </p>\n<p>We know that private data has lower resolution, but in my experiments this should not effect the threshold value… <br>\nAny plan on releasing the hidden test set ? With labels of course.<br>\n<a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <a href=\"https://www.kaggle.com/anjukandru\" target=\"_blank\">@anjukandru</a> <a href=\"https://www.kaggle.com/clairewalsh\" target=\"_blank\">@clairewalsh</a></p>\n<p>Either private data has an issue or my models are really bad 🙃</p>",
  "messages": [
    {
      "id": "2673279",
      "postDate": "02/28/2024 15:52:11",
      "content": "<p>I'm a bit late to the party, but I decided to rerun my model with different thresholds for converting probabilities to masks, to assess how sensitive to it private LB is.<br>\nLocally, my optimal thresholds were between 0.3 and 0.5 for kidneys 1/2/3. Public LB scored quite well with th=0.3 but it's not the case for private … </p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>Private Score</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>th=0.01</td>\n<td>0.733862</td>\n<td>0.819565</td>\n</tr>\n<tr>\n<td>th=0.025</td>\n<td><strong>0.738578</strong></td>\n<td>0.826613</td>\n</tr>\n<tr>\n<td>th=0.05</td>\n<td>0.71296</td>\n<td>0.828603</td>\n</tr>\n<tr>\n<td>th=0.1</td>\n<td>0.678819</td>\n<td>0.829803</td>\n</tr>\n<tr>\n<td>th=0.15</td>\n<td>0.658714</td>\n<td>0.830981</td>\n</tr>\n<tr>\n<td>th=0.2</td>\n<td>0.643396</td>\n<td><strong>0.831588</strong></td>\n</tr>\n<tr>\n<td>th=0.25</td>\n<td>0.629126</td>\n<td>0.829168</td>\n</tr>\n<tr>\n<td>th=0.3</td>\n<td>0.616352</td>\n<td>0.826501</td>\n</tr>\n<tr>\n<td>th=0.5</td>\n<td>0.567844</td>\n<td>0.806398</td>\n</tr>\n</tbody>\n</table>\n<p>The table speaks for itself. This model ranks #3 with an absurdly low threshold of 0.025. </p>\n<p>We know that private data has lower resolution, but in my experiments this should not effect the threshold value… <br>\nAny plan on releasing the hidden test set ? With labels of course.<br>\n<a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <a href=\"https://www.kaggle.com/anjukandru\" target=\"_blank\">@anjukandru</a> <a href=\"https://www.kaggle.com/clairewalsh\" target=\"_blank\">@clairewalsh</a></p>\n<p>Either private data has an issue or my models are really bad 🙃</p>",
      "rawMarkdown": "I'm a bit late to the party, but I decided to rerun my model with different thresholds for converting probabilities to masks, to assess how sensitive to it private LB is.\nLocally, my optimal thresholds were between 0.3 and 0.5 for kidneys 1/2/3. Public LB scored quite well with th=0.3 but it's not the case for private ... \n\n| Threshold | Private Score | Public Score |\n|----------------------------------------|---------------|--------------|\n| th=0.01 | 0.733862      | 0.819565     |\n| th=0.025  | **0.738578**      | 0.826613     |\n| th=0.05   | 0.71296       | 0.828603     |\n| th=0.1  | 0.678819      | 0.829803     |\n| th=0.15   | 0.658714      | 0.830981     |\n| th=0.2   | 0.643396      | **0.831588**     |\n| th=0.25   | 0.629126      | 0.829168     |\n| th=0.3   | 0.616352      | 0.826501     |\n| th=0.5   | 0.567844      | 0.806398     |\n\nThe table speaks for itself. This model ranks #3 with an absurdly low threshold of 0.025. \n\nWe know that private data has lower resolution, but in my experiments this should not effect the threshold value... \nAny plan on releasing the hidden test set ? With labels of course.\n@ryanholbrook @anjukandru @clairewalsh\n\nEither private data has an issue or my models are really bad 🙃",
      "votes": null
    },
    {
      "id": "2675364",
      "postDate": "02/29/2024 21:14:34",
      "content": "<p>I think it has to do with the annotation bias of the data (in particular, how small masks are annotated). </p>\n<p>To illustrate: my optimal thresholds are actually quite low (0.025 is the optimal one) for validation, public, and private when I train with kidney_1 and corresponding pseudo labels </p>\n<p>However, if I start my trainings with kidney_3 (and its corresponding pseudo labels), optimal thresholds converge around 0.3 for validation, public and private. And all of them have bad scores :) </p>\n<p>So overall I would say that it's the issue of FN and you fix that in the simplest way possible - with a low threshold. </p>",
      "rawMarkdown": "I think it has to do with the annotation bias of the data (in particular, how small masks are annotated). \n\nTo illustrate: my optimal thresholds are actually quite low (0.025 is the optimal one) for validation, public, and private when I train with kidney_1 and corresponding pseudo labels \n\nHowever, if I start my trainings with kidney_3 (and its corresponding pseudo labels), optimal thresholds converge around 0.3 for validation, public and private. And all of them have bad scores :) \n\nSo overall I would say that it's the issue of FN and you fix that in the simplest way possible - with a low threshold.",
      "votes": null
    },
    {
      "id": "2677359",
      "postDate": "03/02/2024 06:07:14",
      "content": "<p>I tried to run one of the choosen subs with 0.05, and I got 0.65 private and 0.79 public :D<br>\n(single model, no TTA, had to remove them bcs kernel was crashing for some reasons, maybe scoring function took a little longer than with different threshold, just xy+yz+xz inference)</p>\n<p>For my models and loss function (I chose simple BCE), I had the following optimal thresholds:</p>\n<ol>\n<li>train on kidney_1_dense, validate on kidney_3_dense - optimal threshold was around 0.10 consistently</li>\n<li>train on kidney_1_dense + kidney_3_dense, validation on kidney_2 - optimal threshold ~ 0.30</li>\n</ol>",
      "rawMarkdown": "I tried to run one of the choosen subs with 0.05, and I got 0.65 private and 0.79 public :D\n(single model, no TTA, had to remove them bcs kernel was crashing for some reasons, maybe scoring function took a little longer than with different threshold, just xy+yz+xz inference)\n\nFor my models and loss function (I chose simple BCE), I had the following optimal thresholds:\n1. train on kidney_1_dense, validate on kidney_3_dense - optimal threshold was around 0.10 consistently\n2. train on kidney_1_dense + kidney_3_dense, validation on kidney_2 - optimal threshold ~ 0.30",
      "votes": null
    },
    {
      "id": "2677891",
      "postDate": "03/02/2024 13:29:01",
      "content": "<p>Thanks for your insights !<br>\nSo training on kidney 1 is key, that's interesting.</p>",
      "rawMarkdown": "Thanks for your insights !\nSo training on kidney 1 is key, that's interesting.",
      "votes": null
    },
    {
      "id": "2677960",
      "postDate": "03/02/2024 14:08:08",
      "content": "<p>Yeah. In theory, it shouldn't make any difference. In practice… :) </p>\n<p>That's why I decided not to select the particular kidney and built my 2 finals subs based on kidney_1 and kidney_3 respectively </p>",
      "rawMarkdown": "Yeah. In theory, it shouldn't make any difference. In practice... :) \n\nThat's why I decided not to select the particular kidney and built my 2 finals subs based on kidney_1 and kidney_3 respectively",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2675364,
      "author_name": "ivanpan",
      "author_url": "",
      "post_date": "02/29/2024 21:14:34",
      "content": "<p>I think it has to do with the annotation bias of the data (in particular, how small masks are annotated). </p>\n<p>To illustrate: my optimal thresholds are actually quite low (0.025 is the optimal one) for validation, public, and private when I train with kidney_1 and corresponding pseudo labels </p>\n<p>However, if I start my trainings with kidney_3 (and its corresponding pseudo labels), optimal thresholds converge around 0.3 for validation, public and private. And all of them have bad scores :) </p>\n<p>So overall I would say that it's the issue of FN and you fix that in the simplest way possible - with a low threshold. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2677891,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/02/2024 13:29:01",
          "content": "<p>Thanks for your insights !<br>\nSo training on kidney 1 is key, that's interesting.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2677960,
              "author_name": "ivanpan",
              "author_url": "",
              "post_date": "03/02/2024 14:08:08",
              "content": "<p>Yeah. In theory, it shouldn't make any difference. In practice… :) </p>\n<p>That's why I decided not to select the particular kidney and built my 2 finals subs based on kidney_1 and kidney_3 respectively </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2677359,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "03/02/2024 06:07:14",
      "content": "<p>I tried to run one of the choosen subs with 0.05, and I got 0.65 private and 0.79 public :D<br>\n(single model, no TTA, had to remove them bcs kernel was crashing for some reasons, maybe scoring function took a little longer than with different threshold, just xy+yz+xz inference)</p>\n<p>For my models and loss function (I chose simple BCE), I had the following optimal thresholds:</p>\n<ol>\n<li>train on kidney_1_dense, validate on kidney_3_dense - optimal threshold was around 0.10 consistently</li>\n<li>train on kidney_1_dense + kidney_3_dense, validation on kidney_2 - optimal threshold ~ 0.30</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2673279": "I'm a bit late to the party, but I decided to rerun my model with different thresholds for converting probabilities to masks, to assess how sensitive to it private LB is.\nLocally, my optimal thresholds were between 0.3 and 0.5 for kidneys 1/2/3. Public LB scored quite well with th=0.3 but it's not the case for private ... \n\n| Threshold | Private Score | Public Score |\n|----------------------------------------|---------------|--------------|\n| th=0.01 | 0.733862      | 0.819565     |\n| th=0.025  | **0.738578**      | 0.826613     |\n| th=0.05   | 0.71296       | 0.828603     |\n| th=0.1  | 0.678819      | 0.829803     |\n| th=0.15   | 0.658714      | 0.830981     |\n| th=0.2   | 0.643396      | **0.831588**     |\n| th=0.25   | 0.629126      | 0.829168     |\n| th=0.3   | 0.616352      | 0.826501     |\n| th=0.5   | 0.567844      | 0.806398     |\n\nThe table speaks for itself. This model ranks #3 with an absurdly low threshold of 0.025. \n\nWe know that private data has lower resolution, but in my experiments this should not effect the threshold value... \nAny plan on releasing the hidden test set ? With labels of course.\n@ryanholbrook @anjukandru @clairewalsh\n\nEither private data has an issue or my models are really bad 🙃",
    "2675364": "I think it has to do with the annotation bias of the data (in particular, how small masks are annotated). \n\nTo illustrate: my optimal thresholds are actually quite low (0.025 is the optimal one) for validation, public, and private when I train with kidney_1 and corresponding pseudo labels \n\nHowever, if I start my trainings with kidney_3 (and its corresponding pseudo labels), optimal thresholds converge around 0.3 for validation, public and private. And all of them have bad scores :) \n\nSo overall I would say that it's the issue of FN and you fix that in the simplest way possible - with a low threshold.",
    "2677359": "I tried to run one of the choosen subs with 0.05, and I got 0.65 private and 0.79 public :D\n(single model, no TTA, had to remove them bcs kernel was crashing for some reasons, maybe scoring function took a little longer than with different threshold, just xy+yz+xz inference)\n\nFor my models and loss function (I chose simple BCE), I had the following optimal thresholds:\n1. train on kidney_1_dense, validate on kidney_3_dense - optimal threshold was around 0.10 consistently\n2. train on kidney_1_dense + kidney_3_dense, validation on kidney_2 - optimal threshold ~ 0.30",
    "2677891": "Thanks for your insights !\nSo training on kidney 1 is key, that's interesting.",
    "2677960": "Yeah. In theory, it shouldn't make any difference. In practice... :) \n\nThat's why I decided not to select the particular kidney and built my 2 finals subs based on kidney_1 and kidney_3 respectively"
  },
  "source": "meta"
}