{
  "id": 571858,
  "title": "Rerun Test dataset contains approximately 900 tomograms? ",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/571858",
  "author_name": "",
  "post_date": "2025-04-06T07:16:22.654386700Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Correct me if I'm wrong:</p>\n<pre><code>test/: Directory   directories of dummy test tomograms; the rerun test dataset contains approximately  tomograms. The test data only contain tomograms  one  zero motors.\n</code></pre>\n<p>So in the rerun, there are 900 tomograms — <code>that means 900 unique IDs!</code><br>\nMy model takes 1.55 minutes to predict one, which means 112 seconds per ID.<br>\nSo: <code>112 × 900 = 100800 seconds total,</code> which is around 28 hours to finish all 900 tasks.<br>\nBut our allowed runtime is only 12 hours.</p>\n<p>Question:</p>\n<p>What happens if the model is only able to process, say, 500 IDs? Will scoring be done based on those 500 only, or will the notebook throw an error?</p>\n<p>If we want to calculate the required speed:</p>\n<p>12 hours = <code>12 × 60 × 60 = 43200 seconds</code><br>\nSo to process 900 IDs within that time, we need:<br>\n<code>43200 / 900 = 48 seconds per task</code></p>\n<p>So the model needs to run 1 prediction every 48 seconds.</p>",
  "messages": [
    {
      "id": "3171877",
      "postDate": "04/06/2025 07:16:22",
      "content": "<p>Correct me if I'm wrong:</p>\n<pre><code>test/: Directory   directories of dummy test tomograms; the rerun test dataset contains approximately  tomograms. The test data only contain tomograms  one  zero motors.\n</code></pre>\n<p>So in the rerun, there are 900 tomograms — <code>that means 900 unique IDs!</code><br>\nMy model takes 1.55 minutes to predict one, which means 112 seconds per ID.<br>\nSo: <code>112 × 900 = 100800 seconds total,</code> which is around 28 hours to finish all 900 tasks.<br>\nBut our allowed runtime is only 12 hours.</p>\n<p>Question:</p>\n<p>What happens if the model is only able to process, say, 500 IDs? Will scoring be done based on those 500 only, or will the notebook throw an error?</p>\n<p>If we want to calculate the required speed:</p>\n<p>12 hours = <code>12 × 60 × 60 = 43200 seconds</code><br>\nSo to process 900 IDs within that time, we need:<br>\n<code>43200 / 900 = 48 seconds per task</code></p>\n<p>So the model needs to run 1 prediction every 48 seconds.</p>",
      "rawMarkdown": "Correct me if I'm wrong:\n\n```python\ntest/: Directory with 3 directories of dummy test tomograms; the rerun test dataset contains approximately 900 tomograms. The test data only contain tomograms with one or zero motors.\n```\n\nSo in the rerun, there are 900 tomograms — `that means 900 unique IDs!`\nMy model takes 1.55 minutes to predict one, which means 112 seconds per ID.\nSo: `112 × 900 = 100800 seconds total,` which is around 28 hours to finish all 900 tasks.\nBut our allowed runtime is only 12 hours.\n\nQuestion:\n\nWhat happens if the model is only able to process, say, 500 IDs? Will scoring be done based on those 500 only, or will the notebook throw an error?\n\nIf we want to calculate the required speed:\n\n12 hours = `12 × 60 × 60 = 43200 seconds`\nSo to process 900 IDs within that time, we need:\n`43200 / 900 = 48 seconds per task`\n\nSo the model needs to run 1 prediction every 48 seconds.",
      "votes": null
    },
    {
      "id": "3171880",
      "postDate": "04/06/2025 07:24:40",
      "content": "<p>If you stop prediction after 500 tomograms and mark the last 400 as -1 you will get a score, if not and the notebook does not generate submission.csv after 12 hours you will get an error</p>",
      "rawMarkdown": "If you stop prediction after 500 tomograms and mark the last 400 as -1 you will get a score, if not and the notebook does not generate submission.csv after 12 hours you will get an error",
      "votes": null
    },
    {
      "id": "3171911",
      "postDate": "04/06/2025 08:15:18",
      "content": "<p>That means it's a big problem if our model shows 0.9 accuracy in PBL, yet we still get a very low score in PVL due to the model's speed! This is a challenging task.</p>",
      "rawMarkdown": "That means it's a big problem if our model shows 0.9 accuracy in PBL, yet we still get a very low score in PVL due to the model's speed! This is a challenging task.",
      "votes": null
    },
    {
      "id": "3171928",
      "postDate": "04/06/2025 08:38:37",
      "content": "<p>Sorry, I misunderstood your question. The rerun has no time limit. If your model can complete the public leaderboard part within 12 hours, it will be scored in the private part after the competition ends with no time limit.</p>",
      "rawMarkdown": "Sorry, I misunderstood your question. The rerun has no time limit. If your model can complete the public leaderboard part within 12 hours, it will be scored in the private part after the competition ends with no time limit.",
      "votes": null
    },
    {
      "id": "3172097",
      "postDate": "04/06/2025 13:12:35",
      "content": "<p>More precisely: each submission runs on both the public and the private test set, and must do so within 12 hours. At the end of the competition the score on the private test set is revealed, without rerunning anything.</p>",
      "rawMarkdown": "More precisely: each submission runs on both the public and the private test set, and must do so within 12 hours. At the end of the competition the score on the private test set is revealed, without rerunning anything.",
      "votes": null
    },
    {
      "id": "3172133",
      "postDate": "04/06/2025 14:01:58",
      "content": "<p>I still don't fully understand.</p>\n<p><a href=\"https://www.kaggle.com/fautei\" target=\"_blank\">@fautei</a> said that there is no time limit for the private test set. → <code>True</code><br>\n<a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> said that each submission runs on both the public and private test sets at the same time, and nothing is rerun. The private score is just revealed later.</p>\n<p>So does this mean:<br>\nThe 12-hour submission limit only applies to the public test set (and public leaderboard)?<br>\nBut at the same time, the private test set is also evaluated in the background (with no time limit) correct?&nbsp;</p>",
      "rawMarkdown": "I still don't fully understand.\n\n@fautei said that there is no time limit for the private test set. → `True`\n@jeroencottaar said that each submission runs on both the public and private test sets at the same time, and nothing is rerun. The private score is just revealed later.\n\nSo does this mean:\nThe 12-hour submission limit only applies to the public test set (and public leaderboard)?\nBut at the same time, the private test set is also evaluated in the background (with no time limit) correct?",
      "votes": null
    },
    {
      "id": "3172143",
      "postDate": "04/06/2025 14:07:09",
      "content": "<p>No. When you submit, your code gets a list of 900 tomograms. These are both the public and private test set; your code can't tell which is which. Your code makes a prediction on both. The system then separately scores the public and private set, and reveals only the public score for now.</p>\n<p>So you have 12 hours for the public and the private set, which together comprise about 900 tomograms.</p>",
      "rawMarkdown": "No. When you submit, your code gets a list of 900 tomograms. These are both the public and private test set; your code can't tell which is which. Your code makes a prediction on both. The system then separately scores the public and private set, and reveals only the public score for now.\n\nSo you have 12 hours for the public and the private set, which together comprise about 900 tomograms.",
      "votes": null
    },
    {
      "id": "3172158",
      "postDate": "04/06/2025 14:23:30",
      "content": "<p>Sorry <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>, I bothered you with typing, but could you please type a few more lines?</p>\n<ol>\n<li><code>which together comprise about 900 tomograms.</code> Public set + Private set = 900 tomogram IDs  <code>True</code>?</li>\n<li>Do you know how many tomogram IDs are currently in the public set?</li>\n</ol>",
      "rawMarkdown": "Sorry @jeroencottaar, I bothered you with typing, but could you please type a few more lines?\n1. `which together comprise about 900 tomograms.` Public set + Private set = 900 tomogram IDs  `True`?\n2. Do you know how many tomogram IDs are currently in the public set?",
      "votes": null
    },
    {
      "id": "3172202",
      "postDate": "04/06/2025 15:14:32",
      "content": "<blockquote>\n  <p>Sorry <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>, I bothered you with typing, but could you please type a few more lines?</p>\n  <ol>\n  <li><code>which together comprise about 900 tomograms.</code> Public set + Private set = 900 tomogram IDs  <code>True</code>?</li>\n  </ol>\n</blockquote>\n<p>Correct.</p>\n<blockquote>\n  <ol>\n  <li>Do you know how many tomogram IDs are currently in the public set?</li>\n  </ol>\n</blockquote>\n<p>About 30% are public, so about 270.</p>",
      "rawMarkdown": "> Sorry @jeroencottaar, I bothered you with typing, but could you please type a few more lines?\n> 1. `which together comprise about 900 tomograms.` Public set + Private set = 900 tomogram IDs  `True`?\n\nCorrect.\n\n> 2. Do you know how many tomogram IDs are currently in the public set?\n\nAbout 30% are public, so about 270.",
      "votes": null
    },
    {
      "id": "3173431",
      "postDate": "04/07/2025 22:34:08",
      "content": "<p>Im not sure what is your model but you can run inference on 2xT4 gpus and wrap your model in nn.DataParallel(model). Also you can run in half precision like:</p>\n<pre><code> torch.amp.autocast(device_type=,enabled=,dtype=torch.float16):\n     outputs = model(images)\n</code></pre>\n<p>This might change slightly your results if you didn't train in half precision but T4 gpus are faster in half precision. Cheers!</p>",
      "rawMarkdown": "Im not sure what is your model but you can run inference on 2xT4 gpus and wrap your model in nn.DataParallel(model). Also you can run in half precision like:\n\n```python\nwith torch.amp.autocast(device_type='cuda',enabled=True,dtype=torch.float16):\n     outputs = model(images)\n```\nThis might change slightly your results if you didn't train in half precision but T4 gpus are faster in half precision. Cheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3171880,
      "author_name": "fautei",
      "author_url": "",
      "post_date": "04/06/2025 07:24:40",
      "content": "<p>If you stop prediction after 500 tomograms and mark the last 400 as -1 you will get a score, if not and the notebook does not generate submission.csv after 12 hours you will get an error</p>",
      "votes": null,
      "replies": [
        {
          "id": 3171911,
          "author_name": "sangrampatil5150",
          "author_url": "",
          "post_date": "04/06/2025 08:15:18",
          "content": "<p>That means it's a big problem if our model shows 0.9 accuracy in PBL, yet we still get a very low score in PVL due to the model's speed! This is a challenging task.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3171928,
              "author_name": "fautei",
              "author_url": "",
              "post_date": "04/06/2025 08:38:37",
              "content": "<p>Sorry, I misunderstood your question. The rerun has no time limit. If your model can complete the public leaderboard part within 12 hours, it will be scored in the private part after the competition ends with no time limit.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3172097,
                  "author_name": "jeroencottaar",
                  "author_url": "",
                  "post_date": "04/06/2025 13:12:35",
                  "content": "<p>More precisely: each submission runs on both the public and the private test set, and must do so within 12 hours. At the end of the competition the score on the private test set is revealed, without rerunning anything.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3172133,
                      "author_name": "sangrampatil5150",
                      "author_url": "",
                      "post_date": "04/06/2025 14:01:58",
                      "content": "<p>I still don't fully understand.</p>\n<p><a href=\"https://www.kaggle.com/fautei\" target=\"_blank\">@fautei</a> said that there is no time limit for the private test set. → <code>True</code><br>\n<a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> said that each submission runs on both the public and private test sets at the same time, and nothing is rerun. The private score is just revealed later.</p>\n<p>So does this mean:<br>\nThe 12-hour submission limit only applies to the public test set (and public leaderboard)?<br>\nBut at the same time, the private test set is also evaluated in the background (with no time limit) correct?&nbsp;</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3172143,
                          "author_name": "jeroencottaar",
                          "author_url": "",
                          "post_date": "04/06/2025 14:07:09",
                          "content": "<p>No. When you submit, your code gets a list of 900 tomograms. These are both the public and private test set; your code can't tell which is which. Your code makes a prediction on both. The system then separately scores the public and private set, and reveals only the public score for now.</p>\n<p>So you have 12 hours for the public and the private set, which together comprise about 900 tomograms.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3172158,
                              "author_name": "sangrampatil5150",
                              "author_url": "",
                              "post_date": "04/06/2025 14:23:30",
                              "content": "<p>Sorry <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>, I bothered you with typing, but could you please type a few more lines?</p>\n<ol>\n<li><code>which together comprise about 900 tomograms.</code> Public set + Private set = 900 tomogram IDs  <code>True</code>?</li>\n<li>Do you know how many tomogram IDs are currently in the public set?</li>\n</ol>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3172202,
                                  "author_name": "jeroencottaar",
                                  "author_url": "",
                                  "post_date": "04/06/2025 15:14:32",
                                  "content": "<blockquote>\n  <p>Sorry <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>, I bothered you with typing, but could you please type a few more lines?</p>\n  <ol>\n  <li><code>which together comprise about 900 tomograms.</code> Public set + Private set = 900 tomogram IDs  <code>True</code>?</li>\n  </ol>\n</blockquote>\n<p>Correct.</p>\n<blockquote>\n  <ol>\n  <li>Do you know how many tomogram IDs are currently in the public set?</li>\n  </ol>\n</blockquote>\n<p>About 30% are public, so about 270.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3173431,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "04/07/2025 22:34:08",
      "content": "<p>Im not sure what is your model but you can run inference on 2xT4 gpus and wrap your model in nn.DataParallel(model). Also you can run in half precision like:</p>\n<pre><code> torch.amp.autocast(device_type=,enabled=,dtype=torch.float16):\n     outputs = model(images)\n</code></pre>\n<p>This might change slightly your results if you didn't train in half precision but T4 gpus are faster in half precision. Cheers!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3171877": "Correct me if I'm wrong:\n\n```python\ntest/: Directory with 3 directories of dummy test tomograms; the rerun test dataset contains approximately 900 tomograms. The test data only contain tomograms with one or zero motors.\n```\n\nSo in the rerun, there are 900 tomograms — `that means 900 unique IDs!`\nMy model takes 1.55 minutes to predict one, which means 112 seconds per ID.\nSo: `112 × 900 = 100800 seconds total,` which is around 28 hours to finish all 900 tasks.\nBut our allowed runtime is only 12 hours.\n\nQuestion:\n\nWhat happens if the model is only able to process, say, 500 IDs? Will scoring be done based on those 500 only, or will the notebook throw an error?\n\nIf we want to calculate the required speed:\n\n12 hours = `12 × 60 × 60 = 43200 seconds`\nSo to process 900 IDs within that time, we need:\n`43200 / 900 = 48 seconds per task`\n\nSo the model needs to run 1 prediction every 48 seconds.",
    "3171880": "If you stop prediction after 500 tomograms and mark the last 400 as -1 you will get a score, if not and the notebook does not generate submission.csv after 12 hours you will get an error",
    "3171911": "That means it's a big problem if our model shows 0.9 accuracy in PBL, yet we still get a very low score in PVL due to the model's speed! This is a challenging task.",
    "3171928": "Sorry, I misunderstood your question. The rerun has no time limit. If your model can complete the public leaderboard part within 12 hours, it will be scored in the private part after the competition ends with no time limit.",
    "3172097": "More precisely: each submission runs on both the public and the private test set, and must do so within 12 hours. At the end of the competition the score on the private test set is revealed, without rerunning anything.",
    "3172133": "I still don't fully understand.\n\n@fautei said that there is no time limit for the private test set. → `True`\n@jeroencottaar said that each submission runs on both the public and private test sets at the same time, and nothing is rerun. The private score is just revealed later.\n\nSo does this mean:\nThe 12-hour submission limit only applies to the public test set (and public leaderboard)?\nBut at the same time, the private test set is also evaluated in the background (with no time limit) correct?",
    "3172143": "No. When you submit, your code gets a list of 900 tomograms. These are both the public and private test set; your code can't tell which is which. Your code makes a prediction on both. The system then separately scores the public and private set, and reveals only the public score for now.\n\nSo you have 12 hours for the public and the private set, which together comprise about 900 tomograms.",
    "3172158": "Sorry @jeroencottaar, I bothered you with typing, but could you please type a few more lines?\n1. `which together comprise about 900 tomograms.` Public set + Private set = 900 tomogram IDs  `True`?\n2. Do you know how many tomogram IDs are currently in the public set?",
    "3172202": "> Sorry @jeroencottaar, I bothered you with typing, but could you please type a few more lines?\n> 1. `which together comprise about 900 tomograms.` Public set + Private set = 900 tomogram IDs  `True`?\n\nCorrect.\n\n> 2. Do you know how many tomogram IDs are currently in the public set?\n\nAbout 30% are public, so about 270.",
    "3173431": "Im not sure what is your model but you can run inference on 2xT4 gpus and wrap your model in nn.DataParallel(model). Also you can run in half precision like:\n\n```python\nwith torch.amp.autocast(device_type='cuda',enabled=True,dtype=torch.float16):\n     outputs = model(images)\n```\nThis might change slightly your results if you didn't train in half precision but T4 gpus are faster in half precision. Cheers!"
  },
  "source": "meta"
}