{
  "id": 516365,
  "title": "Runtime limits for public and private scores",
  "url": "/competitions/uspto-explainable-ai/discussion/516365",
  "author_name": "",
  "post_date": "2024-07-02T09:03:41.502700600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi Kagglers, This is my first competition,and as a Kaggle newbie, I feel like a super nut. I'd appreciate your guidance.</p>\n<p>Below are the statements from Overview and Data.</p>\n<blockquote>\n  <p>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n  <p>This leaderboard is calculated with approximately 28% of the test data. The final results will be based on <strong>the other 72%</strong>, so the final standings may be different.</p>\n</blockquote>\n<p>1) If my interpretation is correct, 28%x2,500 patents is being run for current leaderboard public score, and 72%x2,500 patents  will be run for leaderboard private score, which is the final eligible score for this competition. Can anyone confirm?</p>\n<p>2) If that's true, when I submit my notebook now, it should be scored within (28/72)x9hours ( NOT (28/100)x9hours) in order that the same notebook should be run within 9 hours for the final private score. Can anyone confirm?</p>\n<blockquote>\n  <p>CPU Notebook &lt;= 9 hours run-time<br>\n  GPU Notebook &lt;= 9 hours run-time</p>\n  <p>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel. The time required to start the metric notebook and download the data it uses does not count towards the 60 minutes.</p>\n</blockquote>\n<p>3) Does this 60 minutes limit start counting from the point that the whoosh index has been loaded?</p>",
  "messages": [
    {
      "id": "2900413",
      "postDate": "07/02/2024 09:03:41",
      "content": "<p>Hi Kagglers, This is my first competition,and as a Kaggle newbie, I feel like a super nut. I'd appreciate your guidance.</p>\n<p>Below are the statements from Overview and Data.</p>\n<blockquote>\n  <p>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n  <p>This leaderboard is calculated with approximately 28% of the test data. The final results will be based on <strong>the other 72%</strong>, so the final standings may be different.</p>\n</blockquote>\n<p>1) If my interpretation is correct, 28%x2,500 patents is being run for current leaderboard public score, and 72%x2,500 patents  will be run for leaderboard private score, which is the final eligible score for this competition. Can anyone confirm?</p>\n<p>2) If that's true, when I submit my notebook now, it should be scored within (28/72)x9hours ( NOT (28/100)x9hours) in order that the same notebook should be run within 9 hours for the final private score. Can anyone confirm?</p>\n<blockquote>\n  <p>CPU Notebook &lt;= 9 hours run-time<br>\n  GPU Notebook &lt;= 9 hours run-time</p>\n  <p>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel. The time required to start the metric notebook and download the data it uses does not count towards the 60 minutes.</p>\n</blockquote>\n<p>3) Does this 60 minutes limit start counting from the point that the whoosh index has been loaded?</p>",
      "rawMarkdown": "Hi Kagglers, This is my first competition,and as a Kaggle newbie, I feel like a super nut. I'd appreciate your guidance.\n\nBelow are the statements from Overview and Data.\n\n\n>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\n>This leaderboard is calculated with approximately 28% of the test data. The final results will be based on **the other 72%**, so the final standings may be different.\n\n\n1) If my interpretation is correct, 28%x2,500 patents is being run for current leaderboard public score, and 72%x2,500 patents  will be run for leaderboard private score, which is the final eligible score for this competition. Can anyone confirm?\n\n2) If that's true, when I submit my notebook now, it should be scored within (28/72)x9hours ( NOT (28/100)x9hours) in order that the same notebook should be run within 9 hours for the final private score. Can anyone confirm?\n\n\n>CPU Notebook <= 9 hours run-time\n>GPU Notebook <= 9 hours run-time\n\n>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel. The time required to start the metric notebook and download the data it uses does not count towards the 60 minutes.\n\n\n3) Does this 60 minutes limit start counting from the point that the whoosh index has been loaded?",
      "votes": null
    },
    {
      "id": "2903954",
      "postDate": "07/04/2024 05:38:36",
      "content": "<p>I didn't quite understand the 2nd question.</p>\n<h2>1. <strong>How many patents are in the submission?</strong></h2>\n<p>Yes, the test.csv and sample_submission.csv files contain 2500 patents for which you need to compose 2500 of your requests in submission.csv.</p>\n<h2>2. <strong>Operating time of the notebook</strong></h2>\n<p>You have &lt;=9 hours to compile submission.csv for 2500 patents. After that, within 1 hour, a search for neighbouring patents will be performed.</p>\n<h2>3. <strong>Notebook evaluation</strong></h2>\n<p>The public leaderboard shows the result of 28% (2500 * 0.28) test data, when the competition closes you will be shown 72% (2500 * 0.72) result. (notebooks will not be restarted).</p>\n<h2>4.  <strong>Link to previous discussion</strong></h2>\n<p><a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497547#2775330\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497547#2775330</a></p>",
      "rawMarkdown": "I didn't quite understand the 2nd question.\n## 1. **How many patents are in the submission?**\nYes, the test.csv and sample_submission.csv files contain 2500 patents for which you need to compose 2500 of your requests in submission.csv.\n## 2. **Operating time of the notebook**\nYou have <=9 hours to compile submission.csv for 2500 patents. After that, within 1 hour, a search for neighbouring patents will be performed.\n## 3. **Notebook evaluation**\nThe public leaderboard shows the result of 28% (2500 * 0.28) test data, when the competition closes you will be shown 72% (2500 * 0.72) result. (notebooks will not be restarted).\n## 4.  **Link to previous discussion**\nhttps://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497547#2775330",
      "votes": null
    },
    {
      "id": "2904061",
      "postDate": "07/04/2024 06:53:35",
      "content": "<p><a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> Many thanks! So far, I have submitted the submission.csv containing only the patents listed in the test.csv along with their queries, because I thought that test.csv in my Notebook will be replaced with the set of 2,500 patents owned by the host during the public scoring.</p>\n<p>I have one follow-up question: if each participant chooses their own 2,500 patents, and the final score is also based on the set chosen by each participant, is it guaranteed that the scores can be compared fairly? (Or choosing proper 2,500 patents from nearest_neighbors.csv is part of this competition's requirement?) I thought there were 2,500 patents prepared by the host, and the public score was being calculated based on 28% of them…</p>\n<blockquote>\n  <p>Submission File<br>\n  For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in <strong>test.csv</strong>. Your submission file must include a header and have the following format:</p>\n</blockquote>",
      "rawMarkdown": "qurusx Many thanks! So far, I have submitted the submission.csv containing only the patents listed in the test.csv along with their queries, because I thought that test.csv in my Notebook will be replaced with the set of 2,500 patents owned by the host during the public scoring.\n\nI have one follow-up question: if each participant chooses their own 2,500 patents, and the final score is also based on the set chosen by each participant, is it guaranteed that the scores can be compared fairly? (Or choosing proper 2,500 patents from nearest_neighbors.csv is part of this competition's requirement?) I thought there were 2,500 patents prepared by the host, and the public score was being calculated based on 28% of them...\n\n>Submission File\n>For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in **test.csv**. Your submission file must include a header and have the following format:",
      "votes": null
    },
    {
      "id": "2905669",
      "postDate": "07/05/2024 05:29:51",
      "content": "<p>All data <code>/kaggle/input/uspto-explainable-ai</code> for all participants is the same, unchangeable and set at the beginning of the competition. test.csv['publication_number'] = sample_submission.csv['publication_number'].<br>\nIn test.csv, the neighbours are the 50 neighbours for each patent to be found using your queries.<br>\nJust at submission time, the 10 patents in these csv files are replaced by 2500 unknown patents, which your notebook processes again.</p>",
      "rawMarkdown": "All data ` /kaggle/input/uspto-explainable-ai` for all participants is the same, unchangeable and set at the beginning of the competition. test.csv['publication_number'] = sample_submission.csv['publication_number'].\nIn test.csv, the neighbours are the 50 neighbours for each patent to be found using your queries.\nJust at submission time, the 10 patents in these csv files are replaced by 2500 unknown patents, which your notebook processes again.",
      "votes": null
    },
    {
      "id": "2929247",
      "postDate": "07/20/2024 01:03:50",
      "content": "<p><a href=\"https://www.kaggle.com/gowillgo\" target=\"_blank\">@gowillgo</a> <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> Thank you very much for your questions and answers. Can you confirm whether our notebook will run again after the competition is over? I'm a little worried about that</p>",
      "rawMarkdown": "gowillgo @qurusx Thank you very much for your questions and answers. Can you confirm whether our notebook will run again after the competition is over? I'm a little worried about that",
      "votes": null
    },
    {
      "id": "2929610",
      "postDate": "07/20/2024 08:09:03",
      "content": "<p><a href=\"https://www.kaggle.com/alannikos\" target=\"_blank\">@alannikos</a> Considering the runtime on submission, I figure they run with 100% of the dataset but only show the result with 28% of the dataset. So, I don't think the Notebook will be run again after the deadline. Hope this helps.</p>",
      "rawMarkdown": "alannikos Considering the runtime on submission, I figure they run with 100% of the dataset but only show the result with 28% of the dataset. So, I don't think the Notebook will be run again after the deadline. Hope this helps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2903954,
      "author_name": "qurusx",
      "author_url": "",
      "post_date": "07/04/2024 05:38:36",
      "content": "<p>I didn't quite understand the 2nd question.</p>\n<h2>1. <strong>How many patents are in the submission?</strong></h2>\n<p>Yes, the test.csv and sample_submission.csv files contain 2500 patents for which you need to compose 2500 of your requests in submission.csv.</p>\n<h2>2. <strong>Operating time of the notebook</strong></h2>\n<p>You have &lt;=9 hours to compile submission.csv for 2500 patents. After that, within 1 hour, a search for neighbouring patents will be performed.</p>\n<h2>3. <strong>Notebook evaluation</strong></h2>\n<p>The public leaderboard shows the result of 28% (2500 * 0.28) test data, when the competition closes you will be shown 72% (2500 * 0.72) result. (notebooks will not be restarted).</p>\n<h2>4.  <strong>Link to previous discussion</strong></h2>\n<p><a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497547#2775330\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497547#2775330</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2904061,
          "author_name": "gowillgo",
          "author_url": "",
          "post_date": "07/04/2024 06:53:35",
          "content": "<p><a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> Many thanks! So far, I have submitted the submission.csv containing only the patents listed in the test.csv along with their queries, because I thought that test.csv in my Notebook will be replaced with the set of 2,500 patents owned by the host during the public scoring.</p>\n<p>I have one follow-up question: if each participant chooses their own 2,500 patents, and the final score is also based on the set chosen by each participant, is it guaranteed that the scores can be compared fairly? (Or choosing proper 2,500 patents from nearest_neighbors.csv is part of this competition's requirement?) I thought there were 2,500 patents prepared by the host, and the public score was being calculated based on 28% of them…</p>\n<blockquote>\n  <p>Submission File<br>\n  For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in <strong>test.csv</strong>. Your submission file must include a header and have the following format:</p>\n</blockquote>",
          "votes": null,
          "replies": [
            {
              "id": 2905669,
              "author_name": "qurusx",
              "author_url": "",
              "post_date": "07/05/2024 05:29:51",
              "content": "<p>All data <code>/kaggle/input/uspto-explainable-ai</code> for all participants is the same, unchangeable and set at the beginning of the competition. test.csv['publication_number'] = sample_submission.csv['publication_number'].<br>\nIn test.csv, the neighbours are the 50 neighbours for each patent to be found using your queries.<br>\nJust at submission time, the 10 patents in these csv files are replaced by 2500 unknown patents, which your notebook processes again.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2929247,
                  "author_name": "alannikos",
                  "author_url": "",
                  "post_date": "07/20/2024 01:03:50",
                  "content": "<p><a href=\"https://www.kaggle.com/gowillgo\" target=\"_blank\">@gowillgo</a> <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> Thank you very much for your questions and answers. Can you confirm whether our notebook will run again after the competition is over? I'm a little worried about that</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2929610,
                      "author_name": "gowillgo",
                      "author_url": "",
                      "post_date": "07/20/2024 08:09:03",
                      "content": "<p><a href=\"https://www.kaggle.com/alannikos\" target=\"_blank\">@alannikos</a> Considering the runtime on submission, I figure they run with 100% of the dataset but only show the result with 28% of the dataset. So, I don't think the Notebook will be run again after the deadline. Hope this helps.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2900413": "Hi Kagglers, This is my first competition,and as a Kaggle newbie, I feel like a super nut. I'd appreciate your guidance.\n\nBelow are the statements from Overview and Data.\n\n\n>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\n>This leaderboard is calculated with approximately 28% of the test data. The final results will be based on **the other 72%**, so the final standings may be different.\n\n\n1) If my interpretation is correct, 28%x2,500 patents is being run for current leaderboard public score, and 72%x2,500 patents  will be run for leaderboard private score, which is the final eligible score for this competition. Can anyone confirm?\n\n2) If that's true, when I submit my notebook now, it should be scored within (28/72)x9hours ( NOT (28/100)x9hours) in order that the same notebook should be run within 9 hours for the final private score. Can anyone confirm?\n\n\n>CPU Notebook <= 9 hours run-time\n>GPU Notebook <= 9 hours run-time\n\n>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel. The time required to start the metric notebook and download the data it uses does not count towards the 60 minutes.\n\n\n3) Does this 60 minutes limit start counting from the point that the whoosh index has been loaded?",
    "2903954": "I didn't quite understand the 2nd question.\n## 1. **How many patents are in the submission?**\nYes, the test.csv and sample_submission.csv files contain 2500 patents for which you need to compose 2500 of your requests in submission.csv.\n## 2. **Operating time of the notebook**\nYou have <=9 hours to compile submission.csv for 2500 patents. After that, within 1 hour, a search for neighbouring patents will be performed.\n## 3. **Notebook evaluation**\nThe public leaderboard shows the result of 28% (2500 * 0.28) test data, when the competition closes you will be shown 72% (2500 * 0.72) result. (notebooks will not be restarted).\n## 4.  **Link to previous discussion**\nhttps://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497547#2775330",
    "2904061": "qurusx Many thanks! So far, I have submitted the submission.csv containing only the patents listed in the test.csv along with their queries, because I thought that test.csv in my Notebook will be replaced with the set of 2,500 patents owned by the host during the public scoring.\n\nI have one follow-up question: if each participant chooses their own 2,500 patents, and the final score is also based on the set chosen by each participant, is it guaranteed that the scores can be compared fairly? (Or choosing proper 2,500 patents from nearest_neighbors.csv is part of this competition's requirement?) I thought there were 2,500 patents prepared by the host, and the public score was being calculated based on 28% of them...\n\n>Submission File\n>For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in **test.csv**. Your submission file must include a header and have the following format:",
    "2905669": "All data ` /kaggle/input/uspto-explainable-ai` for all participants is the same, unchangeable and set at the beginning of the competition. test.csv['publication_number'] = sample_submission.csv['publication_number'].\nIn test.csv, the neighbours are the 50 neighbours for each patent to be found using your queries.\nJust at submission time, the 10 patents in these csv files are replaced by 2500 unknown patents, which your notebook processes again.",
    "2929247": "gowillgo @qurusx Thank you very much for your questions and answers. Can you confirm whether our notebook will run again after the competition is over? I'm a little worried about that",
    "2929610": "alannikos Considering the runtime on submission, I figure they run with 100% of the dataset but only show the result with 28% of the dataset. So, I don't think the Notebook will be run again after the deadline. Hope this helps."
  },
  "source": "meta"
}