{
  "id": 538018,
  "title": "Is the entire test set inferred during submission or just 38%? (Concern about time limits)",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/538018",
  "author_name": "",
  "post_date": "2024-10-06T12:57:24.203693900Z",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>I’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score? Or is it just 38% inferred initially, with the rest evaluated later?</p>\n<p>My notebook takes quite a while to process the 38%, so it seems pretty likely that it'll exceed the 9-hour limit when running on the full test set. Should I be concerned about this?</p>\n<p>Thanks in advance for any help 🤗</p>",
  "messages": [
    {
      "id": "3008301",
      "postDate": "10/06/2024 12:57:24",
      "content": "<p>Hey everyone,</p>\n<p>I’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score? Or is it just 38% inferred initially, with the rest evaluated later?</p>\n<p>My notebook takes quite a while to process the 38%, so it seems pretty likely that it'll exceed the 9-hour limit when running on the full test set. Should I be concerned about this?</p>\n<p>Thanks in advance for any help 🤗</p>",
      "rawMarkdown": "Hey everyone,\n\nI’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score? Or is it just 38% inferred initially, with the rest evaluated later?\n\nMy notebook takes quite a while to process the 38%, so it seems pretty likely that it'll exceed the 9-hour limit when running on the full test set. Should I be concerned about this?\n\nThanks in advance for any help 🤗",
      "votes": null
    },
    {
      "id": "3008533",
      "postDate": "10/06/2024 17:31:31",
      "content": "<p>Hi,</p>\n<p>\"I’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score\" - you answered your question here.</p>",
      "rawMarkdown": "Hi,\n\n\"I’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score\" - you answered your question here.",
      "votes": null
    },
    {
      "id": "3008554",
      "postDate": "10/06/2024 18:11:39",
      "content": "<p>Hi, thanks for the response! I just wanted to confirm so I don’t risk selecting the wrong model for the final submission.</p>",
      "rawMarkdown": "Hi, thanks for the response! I just wanted to confirm so I don’t risk selecting the wrong model for the final submission.",
      "votes": null
    },
    {
      "id": "3009069",
      "postDate": "10/07/2024 13:39:03",
      "content": "<p>No problem! Well, there is always a risk to select the wrong model (let's say your current best model according to the public leaderboard overfitted the public first 38%, but would perform poorly on the rest 62%). There are different strategies to select the best submission. Some folks select one submission with the best public score and the second submission - the best according to the \"local\" score (your internal cross validation). </p>",
      "rawMarkdown": "No problem! Well, there is always a risk to select the wrong model (let's say your current best model according to the public leaderboard overfitted the public first 38%, but would perform poorly on the rest 62%). There are different strategies to select the best submission. Some folks select one submission with the best public score and the second submission - the best according to the \"local\" score (your internal cross validation).",
      "votes": null
    },
    {
      "id": "3009188",
      "postDate": "10/07/2024 15:57:44",
      "content": "<p>Yea if you get a public leaderboard score you should be good. Your notebook runs on the entire test set during your submission, and at the end of the competition, nothing is run, but everyone's scores are revealed. </p>\n<p>Correct me if I'm wrong, I'm pretty sure that's how it works. </p>\n<p>Thanks. </p>",
      "rawMarkdown": "Yea if you get a public leaderboard score you should be good. Your notebook runs on the entire test set during your submission, and at the end of the competition, nothing is run, but everyone's scores are revealed. \n\nCorrect me if I'm wrong, I'm pretty sure that's how it works. \n\nThanks.",
      "votes": null
    },
    {
      "id": "3009503",
      "postDate": "10/08/2024 02:43:55",
      "content": "<p>The time consumed for your submission is based on testing the entire dataset, so you don't need to worry.</p>",
      "rawMarkdown": "The time consumed for your submission is based on testing the entire dataset, so you don't need to worry.",
      "votes": null
    },
    {
      "id": "3009549",
      "postDate": "10/08/2024 04:08:00",
      "content": "<p>you are wrong on that part</p>\n<p>when we submit it rn the submission is run on 38% of the test dataset and then private score is revealed based on rest of the test dataset</p>",
      "rawMarkdown": "you are wrong on that part\n\nwhen we submit it rn the submission is run on 38% of the test dataset and then private score is revealed based on rest of the test dataset",
      "votes": null
    },
    {
      "id": "3009562",
      "postDate": "10/08/2024 04:34:17",
      "content": "<p>any tips for improvement?? </p>",
      "rawMarkdown": "any tips for improvement??",
      "votes": null
    },
    {
      "id": "3009563",
      "postDate": "10/08/2024 04:35:55",
      "content": "<p>u got any ideas for improvements?</p>",
      "rawMarkdown": "u got any ideas for improvements?",
      "votes": null
    },
    {
      "id": "3010127",
      "postDate": "10/08/2024 17:21:40",
      "content": "<p>I'm pretty sure your wrong. When the submission is run, it is run on the entire test set. Heres what I think happens.</p>\n<ol>\n<li>Submission happens.</li>\n<li>Submission is scored on full test set.</li>\n<li>Public Leaderboard score is revealed on public test set. </li>\n<li>Private leaderboard score is revealed at the end on competition. </li>\n</ol>\n<p>Someone who knows for sure can double check and validate. </p>\n<p>Thanks.</p>",
      "rawMarkdown": "I'm pretty sure your wrong. When the submission is run, it is run on the entire test set. Heres what I think happens.\n1. Submission happens.\n2. Submission is scored on full test set.\n3. Public Leaderboard score is revealed on public test set. \n4. Private leaderboard score is revealed at the end on competition. \n\nSomeone who knows for sure can double check and validate. \n\nThanks.",
      "votes": null
    },
    {
      "id": "3013115",
      "postDate": "10/09/2024 16:59:12",
      "content": "<p><a href=\"https://www.kaggle.com/bopengiowa\" target=\"_blank\">@bopengiowa</a> is correct.</p>",
      "rawMarkdown": "bopengiowa is correct.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3008533,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "10/06/2024 17:31:31",
      "content": "<p>Hi,</p>\n<p>\"I’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score\" - you answered your question here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3008554,
          "author_name": "enzowalter",
          "author_url": "",
          "post_date": "10/06/2024 18:11:39",
          "content": "<p>Hi, thanks for the response! I just wanted to confirm so I don’t risk selecting the wrong model for the final submission.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3009069,
              "author_name": "sergiosaharovskiy",
              "author_url": "",
              "post_date": "10/07/2024 13:39:03",
              "content": "<p>No problem! Well, there is always a risk to select the wrong model (let's say your current best model according to the public leaderboard overfitted the public first 38%, but would perform poorly on the rest 62%). There are different strategies to select the best submission. Some folks select one submission with the best public score and the second submission - the best according to the \"local\" score (your internal cross validation). </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3009563,
                  "author_name": "vedantsinghthakur",
                  "author_url": "",
                  "post_date": "10/08/2024 04:35:55",
                  "content": "<p>u got any ideas for improvements?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3009188,
      "author_name": "bopengiowa",
      "author_url": "",
      "post_date": "10/07/2024 15:57:44",
      "content": "<p>Yea if you get a public leaderboard score you should be good. Your notebook runs on the entire test set during your submission, and at the end of the competition, nothing is run, but everyone's scores are revealed. </p>\n<p>Correct me if I'm wrong, I'm pretty sure that's how it works. </p>\n<p>Thanks. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3009549,
          "author_name": "vedantsinghthakur",
          "author_url": "",
          "post_date": "10/08/2024 04:08:00",
          "content": "<p>you are wrong on that part</p>\n<p>when we submit it rn the submission is run on 38% of the test dataset and then private score is revealed based on rest of the test dataset</p>",
          "votes": null,
          "replies": [
            {
              "id": 3010127,
              "author_name": "bopengiowa",
              "author_url": "",
              "post_date": "10/08/2024 17:21:40",
              "content": "<p>I'm pretty sure your wrong. When the submission is run, it is run on the entire test set. Heres what I think happens.</p>\n<ol>\n<li>Submission happens.</li>\n<li>Submission is scored on full test set.</li>\n<li>Public Leaderboard score is revealed on public test set. </li>\n<li>Private leaderboard score is revealed at the end on competition. </li>\n</ol>\n<p>Someone who knows for sure can double check and validate. </p>\n<p>Thanks.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3013115,
                  "author_name": "sohier",
                  "author_url": "",
                  "post_date": "10/09/2024 16:59:12",
                  "content": "<p><a href=\"https://www.kaggle.com/bopengiowa\" target=\"_blank\">@bopengiowa</a> is correct.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3009503,
      "author_name": "chengweibai",
      "author_url": "",
      "post_date": "10/08/2024 02:43:55",
      "content": "<p>The time consumed for your submission is based on testing the entire dataset, so you don't need to worry.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3009562,
          "author_name": "vedantsinghthakur",
          "author_url": "",
          "post_date": "10/08/2024 04:34:17",
          "content": "<p>any tips for improvement?? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3008301": "Hey everyone,\n\nI’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score? Or is it just 38% inferred initially, with the rest evaluated later?\n\nMy notebook takes quite a while to process the 38%, so it seems pretty likely that it'll exceed the 9-hour limit when running on the full test set. Should I be concerned about this?\n\nThanks in advance for any help 🤗",
    "3008533": "Hi,\n\n\"I’m wondering if the entire test set is inferred during submission and then only 38% is used for the public score\" - you answered your question here.",
    "3008554": "Hi, thanks for the response! I just wanted to confirm so I don’t risk selecting the wrong model for the final submission.",
    "3009069": "No problem! Well, there is always a risk to select the wrong model (let's say your current best model according to the public leaderboard overfitted the public first 38%, but would perform poorly on the rest 62%). There are different strategies to select the best submission. Some folks select one submission with the best public score and the second submission - the best according to the \"local\" score (your internal cross validation).",
    "3009188": "Yea if you get a public leaderboard score you should be good. Your notebook runs on the entire test set during your submission, and at the end of the competition, nothing is run, but everyone's scores are revealed. \n\nCorrect me if I'm wrong, I'm pretty sure that's how it works. \n\nThanks.",
    "3009503": "The time consumed for your submission is based on testing the entire dataset, so you don't need to worry.",
    "3009549": "you are wrong on that part\n\nwhen we submit it rn the submission is run on 38% of the test dataset and then private score is revealed based on rest of the test dataset",
    "3009562": "any tips for improvement??",
    "3009563": "u got any ideas for improvements?",
    "3010127": "I'm pretty sure your wrong. When the submission is run, it is run on the entire test set. Heres what I think happens.\n1. Submission happens.\n2. Submission is scored on full test set.\n3. Public Leaderboard score is revealed on public test set. \n4. Private leaderboard score is revealed at the end on competition. \n\nSomeone who knows for sure can double check and validate. \n\nThanks.",
    "3013115": "bopengiowa is correct."
  },
  "source": "meta"
}