{
  "id": 549271,
  "title": "About the metric, CV vs LB - Bug Fixed",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/549271",
  "author_name": "Ángel Jacinto Sánchez Ruiz",
  "post_date": "2024-12-01T13:15:24.170000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>From the <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussion <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221</a> and the comment of  <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> \"Yeah, exactly, my point is that radius values are given for 1A and it's easy to use them for metric calculation at 10A by mistake.\"</p>\n<p>We've got two metric fixes already so will probably everything be fine and it's just me but… Why my CV scores with overoptimistic predictions (that's with predictions and GT in 10A and radius in 1A, a radius 10 times bigger than it should be) correlate better with public LB than the correct ones (everything in 1A).</p>\n<p>Public LB and overoptimistic ~.5</p>\n<p>correct scales ~.1</p>\n<p>Inside <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">official metric</a> we can find radius are in 1A:</p>\n<p><code>particle_radius = {\n        'apo-ferritin': 60,\n        'beta-amylase': 65,\n        'beta-galactosidase': 90,\n        'ribosome': 150,\n        'thyroglobulin': 130,\n        'virus-like-particle': 135,\n    }</code></p>\n<p>Then they're scaled by distance_multiplier:</p>\n<p><code>particle_radius = {k: v * distance_multiplier for k, v in particle_radius.items()}</code></p>\n<p>The metric call is:</p>\n<p><code>def score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        distance_multiplier: float,\n        beta: int) -&gt; float:</code></p>\n<p>With I assume distance multiplier .5 and beta 4. I am missing something?</p>\n<p>EDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months.</p>",
  "messages": [
    {
      "id": 3060176,
      "postDate": "2024-12-01T13:15:24.170Z",
      "content": "<p>From the <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussion <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221</a> and the comment of  <a href=\"https://www.kaggle.com/ivanpan\" target=\"_blank\">@ivanpan</a> \"Yeah, exactly, my point is that radius values are given for 1A and it's easy to use them for metric calculation at 10A by mistake.\"</p>\n<p>We've got two metric fixes already so will probably everything be fine and it's just me but… Why my CV scores with overoptimistic predictions (that's with predictions and GT in 10A and radius in 1A, a radius 10 times bigger than it should be) correlate better with public LB than the correct ones (everything in 1A).</p>\n<p>Public LB and overoptimistic ~.5</p>\n<p>correct scales ~.1</p>\n<p>Inside <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">official metric</a> we can find radius are in 1A:</p>\n<p><code>particle_radius = {\n        'apo-ferritin': 60,\n        'beta-amylase': 65,\n        'beta-galactosidase': 90,\n        'ribosome': 150,\n        'thyroglobulin': 130,\n        'virus-like-particle': 135,\n    }</code></p>\n<p>Then they're scaled by distance_multiplier:</p>\n<p><code>particle_radius = {k: v * distance_multiplier for k, v in particle_radius.items()}</code></p>\n<p>The metric call is:</p>\n<p><code>def score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        distance_multiplier: float,\n        beta: int) -&gt; float:</code></p>\n<p>With I assume distance multiplier .5 and beta 4. I am missing something?</p>\n<p>EDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months.</p>",
      "rawMarkdown": "From the @hengck23 discussion https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221 and the comment of  @ivanpan \"Yeah, exactly, my point is that radius values are given for 1A and it's easy to use them for metric calculation at 10A by mistake.\"\n\nWe've got two metric fixes already so will probably everything be fine and it's just me but… Why my CV scores with overoptimistic predictions (that's with predictions and GT in 10A and radius in 1A, a radius 10 times bigger than it should be) correlate better with public LB than the correct ones (everything in 1A).\n\nPublic LB and overoptimistic ~.5\n\ncorrect scales ~.1\n\nInside [official metric](https://www.kaggle.com/code/metric/czi-cryoet-84969) we can find radius are in 1A:\n\n`particle_radius = {\n        'apo-ferritin': 60,\n        'beta-amylase': 65,\n        'beta-galactosidase': 90,\n        'ribosome': 150,\n        'thyroglobulin': 130,\n        'virus-like-particle': 135,\n    }`\n\nThen they're scaled by distance_multiplier:\n\n`particle_radius = {k: v * distance_multiplier for k, v in particle_radius.items()}`\n\nThe metric call is:\n\n`def score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        distance_multiplier: float,\n        beta: int) -> float:`\n\nWith I assume distance multiplier .5 and beta 4. I am missing something?\n\nEDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3060176": "From the @hengck23 discussion https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221 and the comment of  @ivanpan \"Yeah, exactly, my point is that radius values are given for 1A and it's easy to use them for metric calculation at 10A by mistake.\"\n\nWe've got two metric fixes already so will probably everything be fine and it's just me but… Why my CV scores with overoptimistic predictions (that's with predictions and GT in 10A and radius in 1A, a radius 10 times bigger than it should be) correlate better with public LB than the correct ones (everything in 1A).\n\nPublic LB and overoptimistic ~.5\n\ncorrect scales ~.1\n\nInside [official metric](https://www.kaggle.com/code/metric/czi-cryoet-84969) we can find radius are in 1A:\n\n`particle_radius = {\n        'apo-ferritin': 60,\n        'beta-amylase': 65,\n        'beta-galactosidase': 90,\n        'ribosome': 150,\n        'thyroglobulin': 130,\n        'virus-like-particle': 135,\n    }`\n\nThen they're scaled by distance_multiplier:\n\n`particle_radius = {k: v * distance_multiplier for k, v in particle_radius.items()}`\n\nThe metric call is:\n\n`def score(\n        solution: pd.DataFrame,\n        submission: pd.DataFrame,\n        row_id_column_name: str,\n        distance_multiplier: float,\n        beta: int) -> float:`\n\nWith I assume distance multiplier .5 and beta 4. I am missing something?\n\nEDIT: Bug fixed, when collecting GT I've been using \"target\" variable that was not changing in the loop. So I've been collecting virus coordinates 5 times per sample. Now everything runs as espected. I'll update my results as soon as possible. \"Fortunately\" there is still about two months."
  }
}