{
  "id": 545799,
  "title": "Why is it difficult to outscore benchmark results?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/545799",
  "author_name": "",
  "post_date": "2024-11-12T05:40:15.612121500Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hey, there!<br>\nThis is the first time I am seeing that Kagglers scores are significantly lower than the benchmark results. Are we all doing something wrong? What could be the possible reasoning?</p>",
  "messages": [
    {
      "id": "3043131",
      "postDate": "11/12/2024 05:40:15",
      "content": "<p>Hey, there!<br>\nThis is the first time I am seeing that Kagglers scores are significantly lower than the benchmark results. Are we all doing something wrong? What could be the possible reasoning?</p>",
      "rawMarkdown": "Hey, there!\nThis is the first time I am seeing that Kagglers scores are significantly lower than the benchmark results. Are we all doing something wrong? What could be the possible reasoning?",
      "votes": null
    },
    {
      "id": "3043313",
      "postDate": "11/12/2024 10:13:23",
      "content": "<p>May be a scoring bug. May be noise gives a lot of false positives and benchmark is too optimistic.</p>",
      "rawMarkdown": "May be a scoring bug. May be noise gives a lot of false positives and benchmark is too optimistic.",
      "votes": null
    },
    {
      "id": "3043352",
      "postDate": "11/12/2024 11:01:42",
      "content": "<p>There's a pretty good discussion thread on this already with interesting info:</p>\n<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381</a></p>",
      "rawMarkdown": "There's a pretty good discussion thread on this already with interesting info:\n\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381)",
      "votes": null
    },
    {
      "id": "3044319",
      "postDate": "11/13/2024 09:02:13",
      "content": "<ol>\n<li>\"May be a scoring bug.\" </li>\n<li>\"May be noise gives a lot of false positives\"</li>\n</ol>\n<p>you may want to think about this:<br>\nhow to prove (1) or (2) using local validation, lb score and probing?</p>\n<hr>\n<p>how to probe fast?<br>\ne.g. if I have 5 hypotheses, I can apply each to 20% of the test data. then i can get results just by making one submission.</p>",
      "rawMarkdown": "1. \"May be a scoring bug.\" \n2. \"May be noise gives a lot of false positives\"\n\n\nyou may want to think about this:\nhow to prove (1) or (2) using local validation, lb score and probing?\n\n---\n\nhow to probe fast?\ne.g. if I have 5 hypotheses, I can apply each to 20% of the test data. then i can get results just by making one submission.",
      "votes": null
    },
    {
      "id": "3044356",
      "postDate": "11/13/2024 09:58:44",
      "content": "<p>Actually I haben't properly dive into it yet. Just talking after what I've read on kaggle. Also after first contact with data the structures look different enough between them and the background. At least regarding the results obtained by real experienced kagglers.</p>\n<p>If nothing weird is happening. Only I can think into a very hard metric. And then… what about benchmark?</p>\n<p>May be yes. I should start to make my own predictions to get a better understanding of what's happening.</p>",
      "rawMarkdown": "Actually I haben't properly dive into it yet. Just talking after what I've read on kaggle. Also after first contact with data the structures look different enough between them and the background. At least regarding the results obtained by real experienced kagglers.\n\nIf nothing weird is happening. Only I can think into a very hard metric. And then... what about benchmark?\n\nMay be yes. I should start to make my own predictions to get a better understanding of what's happening.",
      "votes": null
    },
    {
      "id": "3044359",
      "postDate": "11/13/2024 10:05:35",
      "content": "<p>in machine learning, the assumption is nothing is random. there is some associated distribution with data (even if this distribution is a uniformly random one).</p>\n<p>so if your local cv and public score are different, this means something is different.<br>\nnow if you suspect train and test data are different, you should generate results that don't use data at all.<br>\n(e.g. random submission, fix grid location, etc)</p>\n<p>if that random submission is the same, you can conclude at least the train and test ground truth labels follow the same distribution.</p>",
      "rawMarkdown": "in machine learning, the assumption is nothing is random. there is some associated distribution with data (even if this distribution is a uniformly random one).\n\nso if your local cv and public score are different, this means something is different.\nnow if you suspect train and test data are different, you should generate results that don't use data at all.\n(e.g. random submission, fix grid location, etc)\n\n\nif that random submission is the same, you can conclude at least the train and test ground truth labels follow the same distribution.",
      "votes": null
    },
    {
      "id": "3044891",
      "postDate": "11/13/2024 23:44:31",
      "content": "<p>We've identified the bug in the evaluation and will update soon.</p>\n<p>Thank you for your patience!<br>\nKyle</p>",
      "rawMarkdown": "We've identified the bug in the evaluation and will update soon.\n\nThank you for your patience!\nKyle",
      "votes": null
    },
    {
      "id": "3045033",
      "postDate": "11/14/2024 05:17:18",
      "content": "<p>How can I replicate the F-B evaluation myself? Since in this context, a particle is considered \"true\" if it lies within a factor of 0.5 of the particle of interest's radius. </p>\n<h3>How can I create custom evaluation function?</h3>\n<p>If it is already implemented then please let me know how can I access it.</p>",
      "rawMarkdown": "How can I replicate the F-B evaluation myself? Since in this context, a particle is considered \"true\" if it lies within a factor of 0.5 of the particle of interest's radius. \n### How can I create custom evaluation function?\n\nIf it is already implemented then please let me know how can I access it.",
      "votes": null
    },
    {
      "id": "3045070",
      "postDate": "11/14/2024 06:08:46",
      "content": "<p>Thanks for clarifying, Kyle! :)</p>",
      "rawMarkdown": "Thanks for clarifying, Kyle! :)",
      "votes": null
    },
    {
      "id": "3045295",
      "postDate": "11/14/2024 11:21:16",
      "content": "<p>Here you go: <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">competition metric</a>.</p>",
      "rawMarkdown": "Here you go: [competition metric](https://www.kaggle.com/code/metric/czi-cryoet-84969).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3043313,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "11/12/2024 10:13:23",
      "content": "<p>May be a scoring bug. May be noise gives a lot of false positives and benchmark is too optimistic.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3044319,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/13/2024 09:02:13",
          "content": "<ol>\n<li>\"May be a scoring bug.\" </li>\n<li>\"May be noise gives a lot of false positives\"</li>\n</ol>\n<p>you may want to think about this:<br>\nhow to prove (1) or (2) using local validation, lb score and probing?</p>\n<hr>\n<p>how to probe fast?<br>\ne.g. if I have 5 hypotheses, I can apply each to 20% of the test data. then i can get results just by making one submission.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3044356,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "11/13/2024 09:58:44",
              "content": "<p>Actually I haben't properly dive into it yet. Just talking after what I've read on kaggle. Also after first contact with data the structures look different enough between them and the background. At least regarding the results obtained by real experienced kagglers.</p>\n<p>If nothing weird is happening. Only I can think into a very hard metric. And then… what about benchmark?</p>\n<p>May be yes. I should start to make my own predictions to get a better understanding of what's happening.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3044359,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "11/13/2024 10:05:35",
                  "content": "<p>in machine learning, the assumption is nothing is random. there is some associated distribution with data (even if this distribution is a uniformly random one).</p>\n<p>so if your local cv and public score are different, this means something is different.<br>\nnow if you suspect train and test data are different, you should generate results that don't use data at all.<br>\n(e.g. random submission, fix grid location, etc)</p>\n<p>if that random submission is the same, you can conclude at least the train and test ground truth labels follow the same distribution.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3043352,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "11/12/2024 11:01:42",
      "content": "<p>There's a pretty good discussion thread on this already with interesting info:</p>\n<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3044891,
      "author_name": "kharrington",
      "author_url": "",
      "post_date": "11/13/2024 23:44:31",
      "content": "<p>We've identified the bug in the evaluation and will update soon.</p>\n<p>Thank you for your patience!<br>\nKyle</p>",
      "votes": null,
      "replies": [
        {
          "id": 3045033,
          "author_name": "vigneshwar472",
          "author_url": "",
          "post_date": "11/14/2024 05:17:18",
          "content": "<p>How can I replicate the F-B evaluation myself? Since in this context, a particle is considered \"true\" if it lies within a factor of 0.5 of the particle of interest's radius. </p>\n<h3>How can I create custom evaluation function?</h3>\n<p>If it is already implemented then please let me know how can I access it.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3045295,
              "author_name": "andreizamfir",
              "author_url": "",
              "post_date": "11/14/2024 11:21:16",
              "content": "<p>Here you go: <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">competition metric</a>.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3045070,
          "author_name": "ahsuna123",
          "author_url": "",
          "post_date": "11/14/2024 06:08:46",
          "content": "<p>Thanks for clarifying, Kyle! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3043131": "Hey, there!\nThis is the first time I am seeing that Kagglers scores are significantly lower than the benchmark results. Are we all doing something wrong? What could be the possible reasoning?",
    "3043313": "May be a scoring bug. May be noise gives a lot of false positives and benchmark is too optimistic.",
    "3043352": "There's a pretty good discussion thread on this already with interesting info:\n\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381)",
    "3044319": "1. \"May be a scoring bug.\" \n2. \"May be noise gives a lot of false positives\"\n\n\nyou may want to think about this:\nhow to prove (1) or (2) using local validation, lb score and probing?\n\n---\n\nhow to probe fast?\ne.g. if I have 5 hypotheses, I can apply each to 20% of the test data. then i can get results just by making one submission.",
    "3044356": "Actually I haben't properly dive into it yet. Just talking after what I've read on kaggle. Also after first contact with data the structures look different enough between them and the background. At least regarding the results obtained by real experienced kagglers.\n\nIf nothing weird is happening. Only I can think into a very hard metric. And then... what about benchmark?\n\nMay be yes. I should start to make my own predictions to get a better understanding of what's happening.",
    "3044359": "in machine learning, the assumption is nothing is random. there is some associated distribution with data (even if this distribution is a uniformly random one).\n\nso if your local cv and public score are different, this means something is different.\nnow if you suspect train and test data are different, you should generate results that don't use data at all.\n(e.g. random submission, fix grid location, etc)\n\n\nif that random submission is the same, you can conclude at least the train and test ground truth labels follow the same distribution.",
    "3044891": "We've identified the bug in the evaluation and will update soon.\n\nThank you for your patience!\nKyle",
    "3045033": "How can I replicate the F-B evaluation myself? Since in this context, a particle is considered \"true\" if it lies within a factor of 0.5 of the particle of interest's radius. \n### How can I create custom evaluation function?\n\nIf it is already implemented then please let me know how can I access it.",
    "3045070": "Thanks for clarifying, Kyle! :)",
    "3045295": "Here you go: [competition metric](https://www.kaggle.com/code/metric/czi-cryoet-84969)."
  },
  "source": "meta"
}