{
  "id": 557506,
  "title": "CV vs LB: precision and recall comparison",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/557506",
  "author_name": "Jeroen Cottaar",
  "post_date": "2025-01-19T16:21:00.395000",
  "votes": 22,
  "comment_count": 15,
  "views": 0,
  "content": "<p>A while back I made a post about differences in performance between training set cross-validation (CV) and leaderboard performance (LB) (<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/554815\" target=\"_blank\">link</a>). I dove into this a bit deeper, and notice that both precision and recall are significantly lower on LB compared to CV. This seems odd.</p>\n<p><strong>Method</strong></p>\n<ul>\n<li>My model is a fairly basic Unet model, similar to the example provided. The submitted model is not tweaked in any way to optimize LB score.</li>\n<li>My cross-validation always has 2 sets out of sample and 5 sets in sample. I do this 3 times to cover 6 sets; one set is never out of sample.</li>\n<li>There is quite some variations between model training runs, but nowhere near as large as the differences below.</li>\n<li>I get precision and recall per particle from the LB by submitting only that particle twice, the second time with all predictions doubled. This halves the precision without affecting recall, meaning that we can deduce both precision and recall from the two scores.</li>\n</ul>\n<p><strong>Results</strong></p>\n<table>\n<thead>\n<tr>\n<th>Particle ID</th>\n<th>Precision CV</th>\n<th>Recall CV</th>\n<th>Score CV</th>\n<th>Precision LB</th>\n<th>Recall LB</th>\n<th>Score LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>beta-galactosidase</td>\n<td>0.195</td>\n<td>0.773</td>\n<td>0.658</td>\n<td>0.143</td>\n<td>0.702</td>\n<td>0.571</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>0.210</td>\n<td>0.901</td>\n<td>0.755</td>\n<td>0.121</td>\n<td>0.802</td>\n<td>0.602</td>\n</tr>\n</tbody>\n</table>\n<p>Other particles have similar score between LB and CV.</p>\n<p><strong>Discussion</strong></p>\n<p>Both precision and recall are much lower on the LB. I kind of expected precision, with the idea that there might be some features appearing in the LB data that the model takes as false positives. But I'm surprised by the recall. If the particles in the LB look similar to our training set, I'd expect the same proportion to be found.</p>\n<p>Anyone observed anything similar, or have some bright ideas?</p>",
  "messages": [
    {
      "id": 3100627,
      "postDate": "2025-01-19T16:21:00.397Z",
      "content": "<p>A while back I made a post about differences in performance between training set cross-validation (CV) and leaderboard performance (LB) (<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/554815\" target=\"_blank\">link</a>). I dove into this a bit deeper, and notice that both precision and recall are significantly lower on LB compared to CV. This seems odd.</p>\n<p><strong>Method</strong></p>\n<ul>\n<li>My model is a fairly basic Unet model, similar to the example provided. The submitted model is not tweaked in any way to optimize LB score.</li>\n<li>My cross-validation always has 2 sets out of sample and 5 sets in sample. I do this 3 times to cover 6 sets; one set is never out of sample.</li>\n<li>There is quite some variations between model training runs, but nowhere near as large as the differences below.</li>\n<li>I get precision and recall per particle from the LB by submitting only that particle twice, the second time with all predictions doubled. This halves the precision without affecting recall, meaning that we can deduce both precision and recall from the two scores.</li>\n</ul>\n<p><strong>Results</strong></p>\n<table>\n<thead>\n<tr>\n<th>Particle ID</th>\n<th>Precision CV</th>\n<th>Recall CV</th>\n<th>Score CV</th>\n<th>Precision LB</th>\n<th>Recall LB</th>\n<th>Score LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>beta-galactosidase</td>\n<td>0.195</td>\n<td>0.773</td>\n<td>0.658</td>\n<td>0.143</td>\n<td>0.702</td>\n<td>0.571</td>\n</tr>\n<tr>\n<td>thyroglobulin</td>\n<td>0.210</td>\n<td>0.901</td>\n<td>0.755</td>\n<td>0.121</td>\n<td>0.802</td>\n<td>0.602</td>\n</tr>\n</tbody>\n</table>\n<p>Other particles have similar score between LB and CV.</p>\n<p><strong>Discussion</strong></p>\n<p>Both precision and recall are much lower on the LB. I kind of expected precision, with the idea that there might be some features appearing in the LB data that the model takes as false positives. But I'm surprised by the recall. If the particles in the LB look similar to our training set, I'd expect the same proportion to be found.</p>\n<p>Anyone observed anything similar, or have some bright ideas?</p>",
      "rawMarkdown": "A while back I made a post about differences in performance between training set cross-validation (CV) and leaderboard performance (LB) ([link](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/554815)). I dove into this a bit deeper, and notice that both precision and recall are significantly lower on LB compared to CV. This seems odd.\n\n**Method**\n\n- My model is a fairly basic Unet model, similar to the example provided. The submitted model is not tweaked in any way to optimize LB score.\n- My cross-validation always has 2 sets out of sample and 5 sets in sample. I do this 3 times to cover 6 sets; one set is never out of sample.\n- There is quite some variations between model training runs, but nowhere near as large as the differences below.\n- I get precision and recall per particle from the LB by submitting only that particle twice, the second time with all predictions doubled. This halves the precision without affecting recall, meaning that we can deduce both precision and recall from the two scores.\n\n**Results**\n\n| Particle ID        | Precision CV | Recall CV | Score CV | Precision LB | Recall LB | Score LB |\n| ------------------ | ------------ | --------- | -------- | ------------ | --------- | -------- |\n| beta-galactosidase | 0.195        | 0.773     | 0.658    | 0.143        | 0.702     | 0.571    |\n| thyroglobulin      | 0.210        | 0.901     | 0.755    | 0.121        | 0.802     | 0.602    |\n\nOther particles have similar score between LB and CV.\n\n**Discussion**\n\nBoth precision and recall are much lower on the LB. I kind of expected precision, with the idea that there might be some features appearing in the LB data that the model takes as false positives. But I'm surprised by the recall. If the particles in the LB look similar to our training set, I'd expect the same proportion to be found.\n\nAnyone observed anything similar, or have some bright ideas?\n\n",
      "votes": 22
    },
    {
      "id": 3100703,
      "postDate": "2025-01-19T18:30:42.223Z",
      "content": "<p>Thanks for sharing! When we submit, we can only know F4-score LB. Could you clarify how to calculate precision LB and recall LB? </p>",
      "rawMarkdown": "Thanks for sharing! When we submit, we can only know F4-score LB. Could you clarify how to calculate precision LB and recall LB? ",
      "votes": 1,
      "replies": [
        {
          "id": 3100718,
          "postDate": "2025-01-19T18:41:30.947Z",
          "content": "<p>Make a second submission where you double each prediction (i.e. you predict two particles at the exact same location every time). This halves the precision while not affecting recall. You can retrieve precision and recall as follows (tested offline):</p>\n<p><code># Precision and recall calculator</code><br>\n<code>s1 = 0.163*7/2 # score for baseline submission</code><br>\n<code>s2 = 0.132*7/2 # score for submission with all predictions doubled</code><br>\n<code>precision = 1/(17/s2-17/s1)</code><br>\n<code>recall = -16/(2/precision-17/s2)</code><br>\n<code>print(precision, recall)</code></p>",
          "rawMarkdown": "Make a second submission where you double each prediction (i.e. you predict two particles at the exact same location every time). This halves the precision while not affecting recall. You can retrieve precision and recall as follows (tested offline):\n\n`# Precision and recall calculator`\n`s1 = 0.163*7/2 # score for baseline submission`\n`s2 = 0.132*7/2 # score for submission with all predictions doubled`\n`precision = 1/(17/s2-17/s1)`\n`recall = -16/(2/precision-17/s2)`\n`print(precision, recall)`",
          "votes": 8,
          "replies": [
            {
              "id": 3100786,
              "postDate": "2025-01-19T22:00:39.213Z",
              "content": "<p>Wow, look at the big brain on <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>!  Nice!</p>",
              "rawMarkdown": "Wow, look at the big brain on @jeroencottaar!  Nice!"
            },
            {
              "id": 3100806,
              "postDate": "2025-01-19T22:59:06.497Z",
              "content": "<p>Thanks! That’s clever! :”)</p>",
              "rawMarkdown": "Thanks! That’s clever! :”)"
            }
          ]
        }
      ]
    },
    {
      "id": 3100951,
      "postDate": "2025-01-20T06:37:24.893Z",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> Are your softmax thresholds all set to 0.5?  If, yes, I agree that's weird.  Also, if you're using connected component analysis, how did you select the minimum component size?</p>",
      "rawMarkdown": "@jeroencottaar Are your softmax thresholds all set to 0.5?  If, yes, I agree that's weird.  Also, if you're using connected component analysis, how did you select the minimum component size?",
      "replies": [
        {
          "id": 3100962,
          "postDate": "2025-01-20T06:55:49.797Z",
          "content": "<p>I use a low softmax threshold to catch most particles (still getting some false negatives, including due to mislocating some particles), and then apply a filter based on the size of the connected component. The threshold for that has been optimized to get the best CV performance. (When I said I didn't optimize to LB, I meant specifically that I didn't do any optimization on the submission itself, i.e. it is optimized to CV but not LB).</p>\n<p>I'll play with the thresholds on LB of course, but since both precision and recall are lower on LB this cannot explain the difference.</p>",
          "rawMarkdown": "I use a low softmax threshold to catch most particles (still getting some false negatives, including due to mislocating some particles), and then apply a filter based on the size of the connected component. The threshold for that has been optimized to get the best CV performance. (When I said I didn't optimize to LB, I meant specifically that I didn't do any optimization on the submission itself, i.e. it is optimized to CV but not LB).\n\nI'll play with the thresholds on LB of course, but since both precision and recall are lower on LB this cannot explain the difference.",
          "replies": [
            {
              "id": 3100981,
              "postDate": "2025-01-20T07:29:21.747Z",
              "content": "<p>Why would you think doing that optimization wouldn't cause you to overfit the training/validation data?  I suspect if you set the softmax threshold to 0.5 and select the size randomly, you'll get much more even…though really bad CV/LB values.</p>",
              "rawMarkdown": "Why would you think doing that optimization wouldn't cause you to overfit the training/validation data?  I suspect if you set the softmax threshold to 0.5 and select the size randomly, you'll get much more even...though really bad CV/LB values."
            },
            {
              "id": 3100995,
              "postDate": "2025-01-20T07:48:11.970Z",
              "content": "<p>Most likely the optimal threshold is indeed different on the LB compared to CV. However, if they otherwise behave similar, it shouldn't be the case that <em>both</em> precision and recall go down.</p>\n<p>Forgetting about F4 score entirely, and just consedering recall. On the training set I find 90% (182 out of 202) of thyroglobulin. If thyroglobulin in the public test set looks similar, I'd expect to still find about 90% - regardless of whatever else is going on there. So it seems that the thryglobulin and beta-galactosidase look somehow different in the public test set.</p>",
              "rawMarkdown": "Most likely the optimal threshold is indeed different on the LB compared to CV. However, if they otherwise behave similar, it shouldn't be the case that *both* precision and recall go down.\n\nForgetting about F4 score entirely, and just consedering recall. On the training set I find 90% (182 out of 202) of thyroglobulin. If thyroglobulin in the public test set looks similar, I'd expect to still find about 90% - regardless of whatever else is going on there. So it seems that the thryglobulin and beta-galactosidase look somehow different in the public test set."
            },
            {
              "id": 3100998,
              "postDate": "2025-01-20T07:52:06.693Z",
              "content": "<p>Okay, I see your point.  Might this maybe be related to the chirality issue that's been discussed in other posts?  I can't remember whether some of the 500 were flipped or not.  Although, I thought that was more of a beta-galactosidase issue than thyroglobulin.</p>",
              "rawMarkdown": "Okay, I see your point.  Might this maybe be related to the chirality issue that's been discussed in other posts?  I can't remember whether some of the 500 were flipped or not.  Although, I thought that was more of a beta-galactosidase issue than thyroglobulin."
            },
            {
              "id": 3101014,
              "postDate": "2025-01-20T08:22:37.997Z",
              "content": "<p>Hmmm…  Simulated data is always correct handedness…  All other data, no guarantees.  Or at least that's how I read it:</p>\n<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744</a></p>\n<p>Not convinced this fully explains it though.  But with that said we only got 7 samples out of 500.  Might almost be weirder if the scores were the same.  😀</p>",
              "rawMarkdown": "Hmmm...  Simulated data is always correct handedness...  All other data, no guarantees.  Or at least that's how I read it:\n\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744)\n\nNot convinced this fully explains it though.  But with that said we only got 7 samples out of 500.  Might almost be weirder if the scores were the same.  😀"
            },
            {
              "id": 3101201,
              "postDate": "2025-01-20T13:11:52.770Z",
              "content": "<p>I include flipping in my training data augmentation, so any specific handedness in the training data shouldn't affect matters (and is not taken advantage of). It could all just be matter of these 7 training samples being lucky of course, but the consistency in the precision and recall reduction over both particles makes it seem unlikely to me. </p>\n<p>Probably not much we can do about it in any case…</p>",
              "rawMarkdown": "I include flipping in my training data augmentation, so any specific handedness in the training data shouldn't affect matters (and is not taken advantage of). It could all just be matter of these 7 training samples being lucky of course, but the consistency in the precision and recall reduction over both particles makes it seem unlikely to me. \n\nProbably not much we can do about it in any case..."
            },
            {
              "id": 3101433,
              "postDate": "2025-01-20T18:54:41.473Z",
              "content": "<p>One thing that probably deserves mentioning is the precision/recall tradeoff that occurs by adjusting the softmax threshold is at the pixel level, not the particle level.  For instance, it's relatively straightforward to come up with scenarios where decreasing the pixel threshold actually harms particle recall.  I wonder if some of that is happening here.</p>",
              "rawMarkdown": "One thing that probably deserves mentioning is the precision/recall tradeoff that occurs by adjusting the softmax threshold is at the pixel level, not the particle level.  For instance, it's relatively straightforward to come up with scenarios where decreasing the pixel threshold actually harms particle recall.  I wonder if some of that is happening here."
            },
            {
              "id": 3101487,
              "postDate": "2025-01-20T20:27:32.690Z",
              "content": "<p>For what it's worth, I don't see much difference between softmax threshold and cluster size threshold, where I'm currently using the latter.</p>",
              "rawMarkdown": "For what it's worth, I don't see much difference between softmax threshold and cluster size threshold, where I'm currently using the latter."
            },
            {
              "id": 3105758,
              "postDate": "2025-01-24T00:49:28.313Z",
              "content": "<p>So…  I do have two examples where optimizing parameters with 7-fold cross validation appears to cause worse LB performance than optimizing over a single fold…  I need to do a little more work to fully verify, but that would tend to support what you're seeing.</p>",
              "rawMarkdown": "So...  I do have two examples where optimizing parameters with 7-fold cross validation appears to cause worse LB performance than optimizing over a single fold...  I need to do a little more work to fully verify, but that would tend to support what you're seeing."
            }
          ]
        }
      ]
    },
    {
      "id": 3100836,
      "postDate": "2025-01-20T01:33:08.567Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3100703,
      "author_name": "Pi",
      "author_url": "",
      "post_date": "2025-01-19T18:30:42.223000",
      "content": "<p>Thanks for sharing! When we submit, we can only know F4-score LB. Could you clarify how to calculate precision LB and recall LB? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3100718,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-01-19T18:41:30.947000",
          "content": "<p>Make a second submission where you double each prediction (i.e. you predict two particles at the exact same location every time). This halves the precision while not affecting recall. You can retrieve precision and recall as follows (tested offline):</p>\n<p><code># Precision and recall calculator</code><br>\n<code>s1 = 0.163*7/2 # score for baseline submission</code><br>\n<code>s2 = 0.132*7/2 # score for submission with all predictions doubled</code><br>\n<code>precision = 1/(17/s2-17/s1)</code><br>\n<code>recall = -16/(2/precision-17/s2)</code><br>\n<code>print(precision, recall)</code></p>",
          "votes": 8,
          "replies": [
            {
              "id": 3100786,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-19T22:00:39.213000",
              "content": "<p>Wow, look at the big brain on <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>!  Nice!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3100806,
              "author_name": "Pi",
              "author_url": "",
              "post_date": "2025-01-19T22:59:06.497000",
              "content": "<p>Thanks! That’s clever! :”)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3100951,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2025-01-20T06:37:24.893000",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> Are your softmax thresholds all set to 0.5?  If, yes, I agree that's weird.  Also, if you're using connected component analysis, how did you select the minimum component size?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3100962,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-01-20T06:55:49.797000",
          "content": "<p>I use a low softmax threshold to catch most particles (still getting some false negatives, including due to mislocating some particles), and then apply a filter based on the size of the connected component. The threshold for that has been optimized to get the best CV performance. (When I said I didn't optimize to LB, I meant specifically that I didn't do any optimization on the submission itself, i.e. it is optimized to CV but not LB).</p>\n<p>I'll play with the thresholds on LB of course, but since both precision and recall are lower on LB this cannot explain the difference.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3100981,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-20T07:29:21.747000",
              "content": "<p>Why would you think doing that optimization wouldn't cause you to overfit the training/validation data?  I suspect if you set the softmax threshold to 0.5 and select the size randomly, you'll get much more even…though really bad CV/LB values.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3100995,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-01-20T07:48:11.970000",
              "content": "<p>Most likely the optimal threshold is indeed different on the LB compared to CV. However, if they otherwise behave similar, it shouldn't be the case that <em>both</em> precision and recall go down.</p>\n<p>Forgetting about F4 score entirely, and just consedering recall. On the training set I find 90% (182 out of 202) of thyroglobulin. If thyroglobulin in the public test set looks similar, I'd expect to still find about 90% - regardless of whatever else is going on there. So it seems that the thryglobulin and beta-galactosidase look somehow different in the public test set.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3100998,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-20T07:52:06.693000",
              "content": "<p>Okay, I see your point.  Might this maybe be related to the chirality issue that's been discussed in other posts?  I can't remember whether some of the 500 were flipped or not.  Although, I thought that was more of a beta-galactosidase issue than thyroglobulin.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3101014,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-20T08:22:37.997000",
              "content": "<p>Hmmm…  Simulated data is always correct handedness…  All other data, no guarantees.  Or at least that's how I read it:</p>\n<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/549744</a></p>\n<p>Not convinced this fully explains it though.  But with that said we only got 7 samples out of 500.  Might almost be weirder if the scores were the same.  😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3101201,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-01-20T13:11:52.770000",
              "content": "<p>I include flipping in my training data augmentation, so any specific handedness in the training data shouldn't affect matters (and is not taken advantage of). It could all just be matter of these 7 training samples being lucky of course, but the consistency in the precision and recall reduction over both particles makes it seem unlikely to me. </p>\n<p>Probably not much we can do about it in any case…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3101433,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-20T18:54:41.473000",
              "content": "<p>One thing that probably deserves mentioning is the precision/recall tradeoff that occurs by adjusting the softmax threshold is at the pixel level, not the particle level.  For instance, it's relatively straightforward to come up with scenarios where decreasing the pixel threshold actually harms particle recall.  I wonder if some of that is happening here.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3101487,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2025-01-20T20:27:32.690000",
              "content": "<p>For what it's worth, I don't see much difference between softmax threshold and cluster size threshold, where I'm currently using the latter.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3105758,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2025-01-24T00:49:28.313000",
              "content": "<p>So…  I do have two examples where optimizing parameters with 7-fold cross validation appears to cause worse LB performance than optimizing over a single fold…  I need to do a little more work to fully verify, but that would tend to support what you're seeing.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3100836,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-20T01:33:08.567000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3100627": "A while back I made a post about differences in performance between training set cross-validation (CV) and leaderboard performance (LB) ([link](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/554815)). I dove into this a bit deeper, and notice that both precision and recall are significantly lower on LB compared to CV. This seems odd.\n\n**Method**\n\n- My model is a fairly basic Unet model, similar to the example provided. The submitted model is not tweaked in any way to optimize LB score.\n- My cross-validation always has 2 sets out of sample and 5 sets in sample. I do this 3 times to cover 6 sets; one set is never out of sample.\n- There is quite some variations between model training runs, but nowhere near as large as the differences below.\n- I get precision and recall per particle from the LB by submitting only that particle twice, the second time with all predictions doubled. This halves the precision without affecting recall, meaning that we can deduce both precision and recall from the two scores.\n\n**Results**\n\n| Particle ID        | Precision CV | Recall CV | Score CV | Precision LB | Recall LB | Score LB |\n| ------------------ | ------------ | --------- | -------- | ------------ | --------- | -------- |\n| beta-galactosidase | 0.195        | 0.773     | 0.658    | 0.143        | 0.702     | 0.571    |\n| thyroglobulin      | 0.210        | 0.901     | 0.755    | 0.121        | 0.802     | 0.602    |\n\nOther particles have similar score between LB and CV.\n\n**Discussion**\n\nBoth precision and recall are much lower on the LB. I kind of expected precision, with the idea that there might be some features appearing in the LB data that the model takes as false positives. But I'm surprised by the recall. If the particles in the LB look similar to our training set, I'd expect the same proportion to be found.\n\nAnyone observed anything similar, or have some bright ideas?\n\n",
    "3100703": "Thanks for sharing! When we submit, we can only know F4-score LB. Could you clarify how to calculate precision LB and recall LB? ",
    "3100951": "@jeroencottaar Are your softmax thresholds all set to 0.5?  If, yes, I agree that's weird.  Also, if you're using connected component analysis, how did you select the minimum component size?",
    "3100836": ""
  }
}