{
  "id": 497547,
  "title": "Code Requirements < 9h with 1hr search index? [Solved]",
  "url": "/competitions/uspto-explainable-ai/discussion/497547",
  "author_name": "",
  "post_date": "2024-04-25T02:03:55.577971900Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<blockquote>\n  <p>For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in test.csv. Your submission file must include a header and have the following format:</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>publication_number,query<br>\n  US-2017082634-A1,text AND search<br>\n  US-2017180470-A1,text AND search<br>\n  US-2018029544-A1,text AND search<br>\n  etc.</p>\n</blockquote>\n<ol>\n<li>Queries that <strong>yield fewer than 50 results</strong> will be padded with as many non-matches as necessary to achieve 50 results.</li>\n<li><strong>Queries cannot include more than 50 tokens</strong>, as measured with whoosh_utils.count_query_tokens.</li>\n<li><strong>The metric notebook must finish running your queries in 60 minutes</strong>, not including the time required for loading </li>\n<li>the whoosh index. Note that <strong>the metric uses four Whoosh searchers in parallel</strong>.</li>\n<li>The <strong>number of results for each query is always truncated to 50</strong>.</li>\n</ol>\n<hr>\n<h2></h2>\n<pre><code>                                      \n</code></pre>\n<blockquote>\n  <p><strong>Are we suppose to use only 1 hour search patent index in 9 hours limit for code requirements?</strong></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975. <strong>The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook.</strong> </p>\n</blockquote>\n<h2>We don't have access to train_index so, we can't validate the query?</h2>\n<h2>---</h2>\n<h1>the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.  <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></h1>",
  "messages": [
    {
      "id": "2773965",
      "postDate": "04/25/2024 02:03:55",
      "content": "<blockquote>\n  <p>For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in test.csv. Your submission file must include a header and have the following format:</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>publication_number,query<br>\n  US-2017082634-A1,text AND search<br>\n  US-2017180470-A1,text AND search<br>\n  US-2018029544-A1,text AND search<br>\n  etc.</p>\n</blockquote>\n<ol>\n<li>Queries that <strong>yield fewer than 50 results</strong> will be padded with as many non-matches as necessary to achieve 50 results.</li>\n<li><strong>Queries cannot include more than 50 tokens</strong>, as measured with whoosh_utils.count_query_tokens.</li>\n<li><strong>The metric notebook must finish running your queries in 60 minutes</strong>, not including the time required for loading </li>\n<li>the whoosh index. Note that <strong>the metric uses four Whoosh searchers in parallel</strong>.</li>\n<li>The <strong>number of results for each query is always truncated to 50</strong>.</li>\n</ol>\n<hr>\n<h2></h2>\n<pre><code>                                      \n</code></pre>\n<blockquote>\n  <p><strong>Are we suppose to use only 1 hour search patent index in 9 hours limit for code requirements?</strong></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975. <strong>The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook.</strong> </p>\n</blockquote>\n<h2>We don't have access to train_index so, we can't validate the query?</h2>\n<h2>---</h2>\n<h1>the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.  <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></h1>",
      "rawMarkdown": "> For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in test.csv. Your submission file must include a header and have the following format:\n\n--- \n\n> publication_number,query\nUS-2017082634-A1,text AND search\nUS-2017180470-A1,text AND search\nUS-2018029544-A1,text AND search\netc.\n\n1. Queries that **yield fewer than 50 results** will be padded with as many non-matches as necessary to achieve 50 results.\n2. **Queries cannot include more than 50 tokens**, as measured with whoosh_utils.count_query_tokens.\n3. **The metric notebook must finish running your queries in 60 minutes**, not including the time required for loading \n4. the whoosh index. Note that **the metric uses four Whoosh searchers in parallel**.\n5. The **number of results for each query is always truncated to 50**.\n\n---\n\n## ~~Are we allowed to use search index to check 2,500 queries for 1 hour out of 9 hours?~~\n\n```yaml\npublication number, target publication numbers -> query v1 -> search results v1 < total 1 hour -> new target publications v1 -> improve the query v2 -> search results v2 < total 1 hour -> improve the query v3\n```\n\n> **Are we suppose to use only 1 hour search patent index in 9 hours limit for code requirements?**\n\n---\n\n> train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975. **The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook.** \n\n## We don't have access to train_index so, we can't validate the query?\n\n---\n---\n\n# the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.  @sohier",
      "votes": null
    },
    {
      "id": "2775330",
      "postDate": "04/25/2024 16:12:53",
      "content": "<p>I'm not sure I understand your question but think of it this way: the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.</p>",
      "rawMarkdown": "I'm not sure I understand your question but think of it this way: the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.",
      "votes": null
    },
    {
      "id": "2775333",
      "postDate": "04/25/2024 16:14:08",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, I miss understood it.</p>",
      "rawMarkdown": "Thanks @sohier, I miss understood it.",
      "votes": null
    },
    {
      "id": "2775340",
      "postDate": "04/25/2024 16:19:11",
      "content": "<p>Here is my understanding. </p>\n<p>The one hour requirement is for the metric notebook which is a different notebook from your submission notebook. It is the notebook that calculates the score from your .csv submission output file.  I don't think this changes the code requirements in any way. </p>\n<p>With regard to the index we do have access to a train_index but not the hidden test_index. So I guess this does mean we won't be able to verify our solution using Woosh. </p>",
      "rawMarkdown": "Here is my understanding. \n\nThe one hour requirement is for the metric notebook which is a different notebook from your submission notebook. It is the notebook that calculates the score from your .csv submission output file.  I don't think this changes the code requirements in any way. \n\nWith regard to the index we do have access to a train_index but not the hidden test_index. So I guess this does mean we won't be able to verify our solution using Woosh.",
      "votes": null
    },
    {
      "id": "2775360",
      "postDate": "04/25/2024 16:23:30",
      "content": "<p>Yes its make sense <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a>, as hidden test_index is not available. </p>\n<blockquote>\n  <p>but indirectly you know the 2,500 patents + 50 * 2,500 =&gt; 1,27,500 + maybe extra patents in the hidden test_index ! ( common will make this number less )</p>\n</blockquote>",
      "rawMarkdown": "Yes its make sense @devinanzelmo, as hidden test_index is not available. \n\n> but indirectly you know the 2,500 patents + 50 * 2,500 => 1,27,500 + maybe extra patents in the hidden test_index ! ( common will make this number less )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2775330,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "04/25/2024 16:12:53",
      "content": "<p>I'm not sure I understand your question but think of it this way: the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2775333,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "04/25/2024 16:14:08",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, I miss understood it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2775340,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "04/25/2024 16:19:11",
      "content": "<p>Here is my understanding. </p>\n<p>The one hour requirement is for the metric notebook which is a different notebook from your submission notebook. It is the notebook that calculates the score from your .csv submission output file.  I don't think this changes the code requirements in any way. </p>\n<p>With regard to the index we do have access to a train_index but not the hidden test_index. So I guess this does mean we won't be able to verify our solution using Woosh. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2775360,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "04/25/2024 16:23:30",
          "content": "<p>Yes its make sense <a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a>, as hidden test_index is not available. </p>\n<blockquote>\n  <p>but indirectly you know the 2,500 patents + 50 * 2,500 =&gt; 1,27,500 + maybe extra patents in the hidden test_index ! ( common will make this number less )</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2773965": "> For each publication_number in the test set, you must generate a boolean query that yields all 50 of the target patent IDs specified in test.csv. Your submission file must include a header and have the following format:\n\n--- \n\n> publication_number,query\nUS-2017082634-A1,text AND search\nUS-2017180470-A1,text AND search\nUS-2018029544-A1,text AND search\netc.\n\n1. Queries that **yield fewer than 50 results** will be padded with as many non-matches as necessary to achieve 50 results.\n2. **Queries cannot include more than 50 tokens**, as measured with whoosh_utils.count_query_tokens.\n3. **The metric notebook must finish running your queries in 60 minutes**, not including the time required for loading \n4. the whoosh index. Note that **the metric uses four Whoosh searchers in parallel**.\n5. The **number of results for each query is always truncated to 50**.\n\n---\n\n## ~~Are we allowed to use search index to check 2,500 queries for 1 hour out of 9 hours?~~\n\n```yaml\npublication number, target publication numbers -> query v1 -> search results v1 < total 1 hour -> new target publications v1 -> improve the query v2 -> search results v2 < total 1 hour -> improve the query v3\n```\n\n> **Are we suppose to use only 1 hour search patent index in 9 hours limit for code requirements?**\n\n---\n\n> train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975. **The subset of patents covered by the actual metric index will not be disclosed even to your submission notebook.** \n\n## We don't have access to train_index so, we can't validate the query?\n\n---\n---\n\n# the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.  @sohier",
    "2775330": "I'm not sure I understand your question but think of it this way: the rerun copy of your submission will have 9 hours to generate queries. The metric will then spend up to one hour executing those queries.",
    "2775333": "Thanks @sohier, I miss understood it.",
    "2775340": "Here is my understanding. \n\nThe one hour requirement is for the metric notebook which is a different notebook from your submission notebook. It is the notebook that calculates the score from your .csv submission output file.  I don't think this changes the code requirements in any way. \n\nWith regard to the index we do have access to a train_index but not the hidden test_index. So I guess this does mean we won't be able to verify our solution using Woosh.",
    "2775360": "Yes its make sense @devinanzelmo, as hidden test_index is not available. \n\n> but indirectly you know the 2,500 patents + 50 * 2,500 => 1,27,500 + maybe extra patents in the hidden test_index ! ( common will make this number less )"
  },
  "source": "meta"
}