{
  "id": 571765,
  "title": "submission 0.0000 what is the issue ?",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/571765",
  "author_name": "",
  "post_date": "2025-04-05T12:26:24.453982400Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello<br>\nI submited 2 times and i scored 0.000 twice. Yet i submit with a notebook, it runs first one the fake test dataset (3min and the results are correct) and later on the hidden test dataset which takes 3h to rerun and to score 0.00000. I am sure that a 'submission.csv' is written.<br>\n😃</p>",
  "messages": [
    {
      "id": "3171228",
      "postDate": "04/05/2025 12:26:24",
      "content": "<p>Hello<br>\nI submited 2 times and i scored 0.000 twice. Yet i submit with a notebook, it runs first one the fake test dataset (3min and the results are correct) and later on the hidden test dataset which takes 3h to rerun and to score 0.00000. I am sure that a 'submission.csv' is written.<br>\n😃</p>",
      "rawMarkdown": "Hello\nI submited 2 times and i scored 0.000 twice. Yet i submit with a notebook, it runs first one the fake test dataset (3min and the results are correct) and later on the hidden test dataset which takes 3h to rerun and to score 0.00000. I am sure that a 'submission.csv' is written.\n😃",
      "votes": null
    },
    {
      "id": "3171233",
      "postDate": "04/05/2025 12:43:03",
      "content": "<p>the hidden test set contains many more images than the public one. I had the same issues when I used more threads and a larger batch size than the host example specifies</p>",
      "rawMarkdown": "the hidden test set contains many more images than the public one. I had the same issues when I used more threads and a larger batch size than the host example specifies",
      "votes": null
    },
    {
      "id": "3171242",
      "postDate": "04/05/2025 12:55:03",
      "content": "<p>you mean that i get an error while it's reruning ?</p>",
      "rawMarkdown": "you mean that i get an error while it's reruning ?",
      "votes": null
    },
    {
      "id": "3171249",
      "postDate": "04/05/2025 13:07:16",
      "content": "<p>Yes, or you model can't find any motor.</p>",
      "rawMarkdown": "Yes, or you model can't find any motor.",
      "votes": null
    },
    {
      "id": "3171260",
      "postDate": "04/05/2025 13:26:43",
      "content": "<p>a submission csv with only -1 wouldn't score 0.00 according to the metric right ?</p>",
      "rawMarkdown": "a submission csv with only -1 wouldn't score 0.00 according to the metric right ?",
      "votes": null
    },
    {
      "id": "3171280",
      "postDate": "04/05/2025 14:01:34",
      "content": "<p>I thought so at first, but it seems the metric is only calculated for the motor class, not the background. And in that case, if we only predict -1, the precision in the numerator will be zero, and the whole metric will be 0.</p>",
      "rawMarkdown": "I thought so at first, but it seems the metric is only calculated for the motor class, not the background. And in that case, if we only predict -1, the precision in the numerator will be zero, and the whole metric will be 0.",
      "votes": null
    },
    {
      "id": "3171287",
      "postDate": "04/05/2025 14:12:37",
      "content": "<p>Technically, recall would be 0 too, as you get true positives as 0, since you never predict any positives cases.</p>",
      "rawMarkdown": "Technically, recall would be 0 too, as you get true positives as 0, since you never predict any positives cases.",
      "votes": null
    },
    {
      "id": "3173456",
      "postDate": "04/08/2025 00:09:26",
      "content": "<p><a href=\"https://www.kaggle.com/victorvannobel\" target=\"_blank\">@victorvannobel</a>  Share my experience. I use wrong image size normalization method when doing my GNN approach.<br>\nFor example:</p>\n<pre><code> ():\n     ():\n        ().__init__()\n        .data_path = data_path\n        .data_names = (os.listdir(data_path))[::DATA_CONFIG.jump_step]\n\n     ():\n         (.data_names)\n\n     ():\n        data_name = .data_names[idx]\n        data_path = os.path.join(.data_path, data_name)\n        patch, kp, img_size = point_feature_extractor(data_path)\n        H, W = img_size\n        edge_index = graph_construction(kp)\n        \n        kp[:, ] = kp[:, ] * (H/) \n        kp[:, ] = kp[:, ] * (W/)\n        \n        \n        \n        patch = torch.tensor(patch, dtype=torch.float32)[:, , ...]\n        edge_index = torch.tensor(edge_index)\n        kp = torch.tensor(kp, dtype=torch.int32)\n         Data(x=patch, edge_index=edge_index, y=kp)\n</code></pre>\n<p><code>incorrect normalization</code> will cost you several hours and gets 0.000 score.</p>",
      "rawMarkdown": "victorvannobel  Share my experience. I use wrong image size normalization method when doing my GNN approach.\nFor example:\n\n```python\nclass TomoGraphDataset(Dataset):\n    def __init__(self, data_path):\n        super().__init__()\n        self.data_path = data_path\n        self.data_names = sorted(os.listdir(data_path))[::DATA_CONFIG.jump_step]\n        \n    def len(self):\n        return len(self.data_names)\n\n    def get(self, idx):\n        data_name = self.data_names[idx]\n        data_path = os.path.join(self.data_path, data_name)\n        patch, kp, img_size = point_feature_extractor(data_path)\n        H, W = img_size\n        edge_index = graph_construction(kp)\n        #correct normalization\n        kp[:, 1] = kp[:, 1] * (H/640) \n        kp[:, 0] = kp[:, 0] * (W/640)\n        #incorrect normalization\n        #kp[:, 0] = kp[:, 0] * (H/640) \n        #kp[:, 1] = kp[:, 1] * (W/640)\n        patch = torch.tensor(patch, dtype=torch.float32)[:, None, ...]\n        edge_index = torch.tensor(edge_index)\n        kp = torch.tensor(kp, dtype=torch.int32)\n        return Data(x=patch, edge_index=edge_index, y=kp)\n```\n\n`incorrect normalization` will cost you several hours and gets 0.000 score.",
      "votes": null
    },
    {
      "id": "3173813",
      "postDate": "04/08/2025 11:22:57",
      "content": "<p>If you get a score that means your notebook did write a submission file, otherwise you would get an error. If it's 0 then you're probably predicting no motor for every sample or predicting very inaccurate locations.</p>",
      "rawMarkdown": "If you get a score that means your notebook did write a submission file, otherwise you would get an error. If it's 0 then you're probably predicting no motor for every sample or predicting very inaccurate locations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3171233,
      "author_name": "fautei",
      "author_url": "",
      "post_date": "04/05/2025 12:43:03",
      "content": "<p>the hidden test set contains many more images than the public one. I had the same issues when I used more threads and a larger batch size than the host example specifies</p>",
      "votes": null,
      "replies": [
        {
          "id": 3171242,
          "author_name": "victorvannobel",
          "author_url": "",
          "post_date": "04/05/2025 12:55:03",
          "content": "<p>you mean that i get an error while it's reruning ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3171249,
              "author_name": "fautei",
              "author_url": "",
              "post_date": "04/05/2025 13:07:16",
              "content": "<p>Yes, or you model can't find any motor.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3171260,
                  "author_name": "victorvannobel",
                  "author_url": "",
                  "post_date": "04/05/2025 13:26:43",
                  "content": "<p>a submission csv with only -1 wouldn't score 0.00 according to the metric right ?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3171280,
                      "author_name": "fautei",
                      "author_url": "",
                      "post_date": "04/05/2025 14:01:34",
                      "content": "<p>I thought so at first, but it seems the metric is only calculated for the motor class, not the background. And in that case, if we only predict -1, the precision in the numerator will be zero, and the whole metric will be 0.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3171287,
                          "author_name": "andreizamfir",
                          "author_url": "",
                          "post_date": "04/05/2025 14:12:37",
                          "content": "<p>Technically, recall would be 0 too, as you get true positives as 0, since you never predict any positives cases.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3173456,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "04/08/2025 00:09:26",
      "content": "<p><a href=\"https://www.kaggle.com/victorvannobel\" target=\"_blank\">@victorvannobel</a>  Share my experience. I use wrong image size normalization method when doing my GNN approach.<br>\nFor example:</p>\n<pre><code> ():\n     ():\n        ().__init__()\n        .data_path = data_path\n        .data_names = (os.listdir(data_path))[::DATA_CONFIG.jump_step]\n\n     ():\n         (.data_names)\n\n     ():\n        data_name = .data_names[idx]\n        data_path = os.path.join(.data_path, data_name)\n        patch, kp, img_size = point_feature_extractor(data_path)\n        H, W = img_size\n        edge_index = graph_construction(kp)\n        \n        kp[:, ] = kp[:, ] * (H/) \n        kp[:, ] = kp[:, ] * (W/)\n        \n        \n        \n        patch = torch.tensor(patch, dtype=torch.float32)[:, , ...]\n        edge_index = torch.tensor(edge_index)\n        kp = torch.tensor(kp, dtype=torch.int32)\n         Data(x=patch, edge_index=edge_index, y=kp)\n</code></pre>\n<p><code>incorrect normalization</code> will cost you several hours and gets 0.000 score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3173813,
      "author_name": "tennogh",
      "author_url": "",
      "post_date": "04/08/2025 11:22:57",
      "content": "<p>If you get a score that means your notebook did write a submission file, otherwise you would get an error. If it's 0 then you're probably predicting no motor for every sample or predicting very inaccurate locations.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3171228": "Hello\nI submited 2 times and i scored 0.000 twice. Yet i submit with a notebook, it runs first one the fake test dataset (3min and the results are correct) and later on the hidden test dataset which takes 3h to rerun and to score 0.00000. I am sure that a 'submission.csv' is written.\n😃",
    "3171233": "the hidden test set contains many more images than the public one. I had the same issues when I used more threads and a larger batch size than the host example specifies",
    "3171242": "you mean that i get an error while it's reruning ?",
    "3171249": "Yes, or you model can't find any motor.",
    "3171260": "a submission csv with only -1 wouldn't score 0.00 according to the metric right ?",
    "3171280": "I thought so at first, but it seems the metric is only calculated for the motor class, not the background. And in that case, if we only predict -1, the precision in the numerator will be zero, and the whole metric will be 0.",
    "3171287": "Technically, recall would be 0 too, as you get true positives as 0, since you never predict any positives cases.",
    "3173456": "victorvannobel  Share my experience. I use wrong image size normalization method when doing my GNN approach.\nFor example:\n\n```python\nclass TomoGraphDataset(Dataset):\n    def __init__(self, data_path):\n        super().__init__()\n        self.data_path = data_path\n        self.data_names = sorted(os.listdir(data_path))[::DATA_CONFIG.jump_step]\n        \n    def len(self):\n        return len(self.data_names)\n\n    def get(self, idx):\n        data_name = self.data_names[idx]\n        data_path = os.path.join(self.data_path, data_name)\n        patch, kp, img_size = point_feature_extractor(data_path)\n        H, W = img_size\n        edge_index = graph_construction(kp)\n        #correct normalization\n        kp[:, 1] = kp[:, 1] * (H/640) \n        kp[:, 0] = kp[:, 0] * (W/640)\n        #incorrect normalization\n        #kp[:, 0] = kp[:, 0] * (H/640) \n        #kp[:, 1] = kp[:, 1] * (W/640)\n        patch = torch.tensor(patch, dtype=torch.float32)[:, None, ...]\n        edge_index = torch.tensor(edge_index)\n        kp = torch.tensor(kp, dtype=torch.int32)\n        return Data(x=patch, edge_index=edge_index, y=kp)\n```\n\n`incorrect normalization` will cost you several hours and gets 0.000 score.",
    "3173813": "If you get a score that means your notebook did write a submission file, otherwise you would get an error. If it's 0 then you're probably predicting no motor for every sample or predicting very inaccurate locations."
  },
  "source": "meta"
}