{
  "id": 516104,
  "title": "Tips for advanced usage of Whoosh",
  "url": "/competitions/uspto-explainable-ai/discussion/516104",
  "author_name": "",
  "post_date": "2024-07-01T10:42:06.046861600Z",
  "votes": 44,
  "comment_count": 9,
  "views": 0,
  "content": "<p>While researching whoosh, I discovered some features that can improve the LB score. Since rankings based on this knowledge seem to deviate from the competition's purpose, I’m sharing these insights.</p>\n<h1>1. Tips for saving tokens</h1>\n<h2>1.1. Omit AND token</h2>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/querylang.html#boolean-operators\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/querylang.html#boolean-operators</a></p>\n<ul>\n<li><code>ti:dog ti:cat</code> is equivalent to <code>ti:dog AND ti:cat</code></li>\n<li><code>ti:dog ab:cat clm:fox</code> is equivalent to <code>ti:dog AND ab:cat AND clm:fox</code></li>\n</ul>\n<h2>1.2. Omit ADJ token</h2>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/querylang.html#individual-terms-and-phrases\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/querylang.html#individual-terms-and-phrases</a></p>\n<ul>\n<li><code>ti:\"open sesame\"</code> is equivalent to <code>ti:open ADJ1 ti:sesame</code></li>\n<li><code>ti:\"open sesame\"~2</code> is equivalent to <code>ti:open ADJ2 ti:sesame</code></li>\n</ul>\n<h2>1.3. Search across fields</h2>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/parsing.html#letting-the-user-search-multiple-fields-by-default\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/parsing.html#letting-the-user-search-multiple-fields-by-default</a></p>\n<ul>\n<li><code>apple</code> is equivalent to <code>ti:apple OR ab:apple OR clm:apple OR detd:apple OR cpc:apple</code></li>\n</ul>\n<h1>2. Tips for result order</h1>\n<ul>\n<li>Results are returned in descending order of their TF-IDF scores</li>\n<li>If the scores are the same, the results are returned in the order they were indexed in whoosh. Therefore, changing the order of the query does not affect the order of the results.</li>\n</ul>\n<h1>Examples</h1>\n<pre><code> whoosh_utils\n\ndocs = [\n    {:, :, :, :, :, : []},\n    {:, :, :, :, :, : []},\n    {:, :, :, :, :, : []},\n]\nwhoosh_utils.create_index(, docs)\n\nsample_index = whoosh_utils.load_index()\nsample_searcher = whoosh_utils.get_searcher(sample_index)\nqp = whoosh_utils.get_query_parser()\nqv = whoosh_utils.QueryValidator()\n\n query  [, , , , , , ]:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    results = whoosh_utils.execute_query(query, qp, sample_searcher)\n    (query, results, num_tokens)\n</code></pre>\n<pre><code>\nti:dog ti:cat [, ] \n\nab: [] \nab:~ [, ] \n\ndog [, , ] \n\nclm:cat [, , ] \ndetd:cat OR detd:dog [, ] \ndetd:dog OR detd:cat [, ] \n</code></pre>\n<p>Hopefully this will help you to focus on the main part of the competition.</p>",
  "messages": [
    {
      "id": "2898783",
      "postDate": "07/01/2024 10:42:06",
      "content": "<p>While researching whoosh, I discovered some features that can improve the LB score. Since rankings based on this knowledge seem to deviate from the competition's purpose, I’m sharing these insights.</p>\n<h1>1. Tips for saving tokens</h1>\n<h2>1.1. Omit AND token</h2>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/querylang.html#boolean-operators\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/querylang.html#boolean-operators</a></p>\n<ul>\n<li><code>ti:dog ti:cat</code> is equivalent to <code>ti:dog AND ti:cat</code></li>\n<li><code>ti:dog ab:cat clm:fox</code> is equivalent to <code>ti:dog AND ab:cat AND clm:fox</code></li>\n</ul>\n<h2>1.2. Omit ADJ token</h2>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/querylang.html#individual-terms-and-phrases\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/querylang.html#individual-terms-and-phrases</a></p>\n<ul>\n<li><code>ti:\"open sesame\"</code> is equivalent to <code>ti:open ADJ1 ti:sesame</code></li>\n<li><code>ti:\"open sesame\"~2</code> is equivalent to <code>ti:open ADJ2 ti:sesame</code></li>\n</ul>\n<h2>1.3. Search across fields</h2>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/parsing.html#letting-the-user-search-multiple-fields-by-default\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/parsing.html#letting-the-user-search-multiple-fields-by-default</a></p>\n<ul>\n<li><code>apple</code> is equivalent to <code>ti:apple OR ab:apple OR clm:apple OR detd:apple OR cpc:apple</code></li>\n</ul>\n<h1>2. Tips for result order</h1>\n<ul>\n<li>Results are returned in descending order of their TF-IDF scores</li>\n<li>If the scores are the same, the results are returned in the order they were indexed in whoosh. Therefore, changing the order of the query does not affect the order of the results.</li>\n</ul>\n<h1>Examples</h1>\n<pre><code> whoosh_utils\n\ndocs = [\n    {:, :, :, :, :, : []},\n    {:, :, :, :, :, : []},\n    {:, :, :, :, :, : []},\n]\nwhoosh_utils.create_index(, docs)\n\nsample_index = whoosh_utils.load_index()\nsample_searcher = whoosh_utils.get_searcher(sample_index)\nqp = whoosh_utils.get_query_parser()\nqv = whoosh_utils.QueryValidator()\n\n query  [, , , , , , ]:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    results = whoosh_utils.execute_query(query, qp, sample_searcher)\n    (query, results, num_tokens)\n</code></pre>\n<pre><code>\nti:dog ti:cat [, ] \n\nab: [] \nab:~ [, ] \n\ndog [, , ] \n\nclm:cat [, , ] \ndetd:cat OR detd:dog [, ] \ndetd:dog OR detd:cat [, ] \n</code></pre>\n<p>Hopefully this will help you to focus on the main part of the competition.</p>",
      "rawMarkdown": "While researching whoosh, I discovered some features that can improve the LB score. Since rankings based on this knowledge seem to deviate from the competition's purpose, I’m sharing these insights.\n\n# 1. Tips for saving tokens\n## 1.1. Omit AND token\nhttps://whoosh.readthedocs.io/en/latest/querylang.html#boolean-operators\n- `ti:dog ti:cat` is equivalent to `ti:dog AND ti:cat`\n- `ti:dog ab:cat clm:fox` is equivalent to `ti:dog AND ab:cat AND clm:fox`\n\n## 1.2. Omit ADJ token\nhttps://whoosh.readthedocs.io/en/latest/querylang.html#individual-terms-and-phrases\n- `ti:\"open sesame\"` is equivalent to `ti:open ADJ1 ti:sesame`\n- `ti:\"open sesame\"~2` is equivalent to `ti:open ADJ2 ti:sesame`\n\n## 1.3. Search across fields\nhttps://whoosh.readthedocs.io/en/latest/parsing.html#letting-the-user-search-multiple-fields-by-default\n- `apple` is equivalent to `ti:apple OR ab:apple OR clm:apple OR detd:apple OR cpc:apple`\n\n# 2. Tips for result order\n- Results are returned in descending order of their TF-IDF scores\n- If the scores are the same, the results are returned in the order they were indexed in whoosh. Therefore, changing the order of the query does not affect the order of the results.\n\n# Examples\n```py\nimport whoosh_utils\n\ndocs = [\n    {\"publication_number\":\"one\", \"title\":\"dog dummy cat\", \"abstract\":\"sesame dummy open\", \"claims\":\"cat\", \"description\":\"dummy\", \"cpc\": [\"dummy\"]},\n    {\"publication_number\":\"two\", \"title\":\"cat dummy dummy dog\", \"abstract\":\"open sesame\", \"claims\":\"cat cat\", \"description\":\"cat\", \"cpc\": [\"dummy\"]},\n    {\"publication_number\":\"three\", \"title\":\"cat\", \"abstract\":\"open dummy dummy sesame\", \"claims\":\"cat cat cat\", \"description\":\"dog\", \"cpc\": [\"dummy\"]},\n]\nwhoosh_utils.create_index(f\"sample\", docs)\n\nsample_index = whoosh_utils.load_index(f\"sample\")\nsample_searcher = whoosh_utils.get_searcher(sample_index)\nqp = whoosh_utils.get_query_parser()\nqv = whoosh_utils.QueryValidator()\n\nfor query in [\"ti:dog ti:cat\", 'ab:\"open sesame\"', 'ab:\"open sesame\"~3', 'dog', 'clm:cat', 'detd:cat OR detd:dog', 'detd:dog OR detd:cat']:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    results = whoosh_utils.execute_query(query, qp, sample_searcher)\n    print(query, results, num_tokens)\n```\n\n```py\n# Example of checking \"1.1. Omit AND token\"\nti:dog ti:cat ['one', 'two'] 2\n# Example of checking \"1.2. Omit ADJ token\"\nab:\"open sesame\" ['two'] 2\nab:\"open sesame\"~3 ['two', 'three'] 2\n# Example of checking \"1.3. Search across fields\"\ndog ['three', 'one', 'two'] 1\n# Example of checking \"2. Tips for result order\"\nclm:cat ['three', 'two', 'one'] 1\ndetd:cat OR detd:dog ['two', 'three'] 3\ndetd:dog OR detd:cat ['two', 'three'] 3\n```\n\nHopefully this will help you to focus on the main part of the competition.",
      "votes": null
    },
    {
      "id": "2898799",
      "postDate": "07/01/2024 10:50:59",
      "content": "<p>Thank you, valuable information!<br>\nAre operators also counted before tokens are counted? (AND, OR, etc.)?</p>",
      "rawMarkdown": "Thank you, valuable information!\nAre operators also counted before tokens are counted? (AND, OR, etc.)?",
      "votes": null
    },
    {
      "id": "2898823",
      "postDate": "07/01/2024 11:01:21",
      "content": "<p>Thanks! How can we retrieve TF-IDF scores of all documents from the index?</p>",
      "rawMarkdown": "Thanks! How can we retrieve TF-IDF scores of all documents from the index?",
      "votes": null
    },
    {
      "id": "2898891",
      "postDate": "07/01/2024 11:36:53",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/overview\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/overview</a></p>\n<blockquote>\n  <p>Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.</p>\n</blockquote>\n<p>Is the question asking if operators are counted as well? If so, yes, operators are a type of token. (In my understanding)</p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/uspto-explainable-ai/overview\n> Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.\n\nIs the question asking if operators are counted as well? If so, yes, operators are a type of token. (In my understanding)",
      "votes": null
    },
    {
      "id": "2898897",
      "postDate": "07/01/2024 11:43:27",
      "content": "<p>You can directly look at the contents of whoosh.searching.Hit</p>\n<pre><code> () -&gt; :\n     (query) &gt; :\n         ValueError()\n       query:\n         ValueError()\n\n    to_search = qp.parse(query)\n    results = searcher.search(to_search, limit=results_limit)\n    \n    results = [(x.score, x[])  x  results]\n     (results) &lt;= results_limit\n     results\n\n...\n\n query  [, , , , , , ]:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    \n    results = execute_query(query, qp, sample_searcher)\n    (query, results, num_tokens)\n</code></pre>\n<pre><code>ti:dog ti:cat [(, ), (, )] \nab: [(, )] \nab:~ [(, ), (, )] \ndog [(, ), (, ), (, )] \nclm:cat [(, ), (, ), (, )] \ndetd:cat OR detd:dog [(, ), (, )] \ndetd:dog OR detd:cat [(, ), (, )] \n</code></pre>\n<p>For more details: <a href=\"https://whoosh.readthedocs.io/en/latest/api/searching.html#whoosh.searching.Hit\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/api/searching.html#whoosh.searching.Hit</a></p>",
      "rawMarkdown": "You can directly look at the contents of whoosh.searching.Hit\n\n```py\ndef execute_query(query: str, qp, searcher, results_limit=50) -> list:\n    if len(query) > 10_000:\n        raise ValueError('Query length at exceeds 10,000 characters.')\n    if 'id:' in query:\n        raise ValueError('Searching for specific patent IDs is banned.')\n\n    to_search = qp.parse(query)\n    results = searcher.search(to_search, limit=results_limit)\n    # results = [x['id'] for x in results]\n    results = [(x.score, x['id']) for x in results]\n    assert len(results) <= results_limit\n    return results\n\n...\n\nfor query in [\"ti:dog ti:cat\", 'ab:\"open sesame\"', 'ab:\"open sesame\"~3', 'dog', 'clm:cat', 'detd:cat OR detd:dog', 'detd:dog OR detd:cat']:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    # results = whoosh_utils.execute_query(query, qp, sample_searcher)\n    results = execute_query(query, qp, sample_searcher)\n    print(query, results, num_tokens)\n```\n\n```py\nti:dog ti:cat [(1.7123179275482192, 'one'), (1.7123179275482192, 'two')] 2\nab:\"open sesame\" [(1.4246358550964382, 'two')] 2\nab:\"open sesame\"~3 [(1.4246358550964382, 'two'), (1.4246358550964382, 'three')] 2\ndog [(1.4054651081081644, 'three'), (1.0, 'one'), (1.0, 'two')] 1\nclm:cat [(2.136953782644657, 'three'), (1.4246358550964382, 'two'), (0.7123179275482191, 'one')] 1\ndetd:cat OR detd:dog [(1.4054651081081644, 'two'), (1.4054651081081644, 'three')] 3\ndetd:dog OR detd:cat [(1.4054651081081644, 'two'), (1.4054651081081644, 'three')] 3\n```\n\nFor more details: https://whoosh.readthedocs.io/en/latest/api/searching.html#whoosh.searching.Hit",
      "votes": null
    },
    {
      "id": "2899071",
      "postDate": "07/01/2024 13:31:11",
      "content": "<p>Here is another one:<br>\nNOT cpc:*  - this query returns patents with no CPC codes linked to them</p>",
      "rawMarkdown": "Here is another one:\nNOT cpc:*  - this query returns patents with no CPC codes linked to them",
      "votes": null
    },
    {
      "id": "2902831",
      "postDate": "07/03/2024 13:28:10",
      "content": "<p>Thanks ! But whoosh runs indefinitely when I add this … Do you know why ?</p>",
      "rawMarkdown": "Thanks ! But whoosh runs indefinitely when I add this ... Do you know why ?",
      "votes": null
    },
    {
      "id": "2902845",
      "postDate": "07/03/2024 13:38:44",
      "content": "<p>Using wildcards (* and ?) makes queries run longer. I'm pretty sure that if every query contains a wildcard, then the 2500 queries will take more than the allowed hour. </p>\n<p>When I tested this specific subquery \"NOT cpc:*\", the query containing it would take about 15 seconds to run.</p>",
      "rawMarkdown": "Using wildcards (* and ?) makes queries run longer. I'm pretty sure that if every query contains a wildcard, then the 2500 queries will take more than the allowed hour. \n\nWhen I tested this specific subquery \"NOT cpc:*\", the query containing it would take about 15 seconds to run.",
      "votes": null
    },
    {
      "id": "2913873",
      "postDate": "07/09/2024 16:50:57",
      "content": "<p>Excellent work</p>",
      "rawMarkdown": "Excellent work",
      "votes": null
    },
    {
      "id": "2920931",
      "postDate": "07/13/2024 23:14:56",
      "content": "<p>Nice work. Thanks</p>",
      "rawMarkdown": "Nice work. Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2898799,
      "author_name": "qurusx",
      "author_url": "",
      "post_date": "07/01/2024 10:50:59",
      "content": "<p>Thank you, valuable information!<br>\nAre operators also counted before tokens are counted? (AND, OR, etc.)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2898891,
          "author_name": "sash2104",
          "author_url": "",
          "post_date": "07/01/2024 11:36:53",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/overview\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/overview</a></p>\n<blockquote>\n  <p>Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.</p>\n</blockquote>\n<p>Is the question asking if operators are counted as well? If so, yes, operators are a type of token. (In my understanding)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2898823,
      "author_name": "tangtunyu",
      "author_url": "",
      "post_date": "07/01/2024 11:01:21",
      "content": "<p>Thanks! How can we retrieve TF-IDF scores of all documents from the index?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2898897,
          "author_name": "sash2104",
          "author_url": "",
          "post_date": "07/01/2024 11:43:27",
          "content": "<p>You can directly look at the contents of whoosh.searching.Hit</p>\n<pre><code> () -&gt; :\n     (query) &gt; :\n         ValueError()\n       query:\n         ValueError()\n\n    to_search = qp.parse(query)\n    results = searcher.search(to_search, limit=results_limit)\n    \n    results = [(x.score, x[])  x  results]\n     (results) &lt;= results_limit\n     results\n\n...\n\n query  [, , , , , , ]:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    \n    results = execute_query(query, qp, sample_searcher)\n    (query, results, num_tokens)\n</code></pre>\n<pre><code>ti:dog ti:cat [(, ), (, )] \nab: [(, )] \nab:~ [(, ), (, )] \ndog [(, ), (, ), (, )] \nclm:cat [(, ), (, ), (, )] \ndetd:cat OR detd:dog [(, ), (, )] \ndetd:dog OR detd:cat [(, ), (, )] \n</code></pre>\n<p>For more details: <a href=\"https://whoosh.readthedocs.io/en/latest/api/searching.html#whoosh.searching.Hit\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/api/searching.html#whoosh.searching.Hit</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2899071,
      "author_name": "apparition",
      "author_url": "",
      "post_date": "07/01/2024 13:31:11",
      "content": "<p>Here is another one:<br>\nNOT cpc:*  - this query returns patents with no CPC codes linked to them</p>",
      "votes": null,
      "replies": [
        {
          "id": 2902831,
          "author_name": "vincentschuler",
          "author_url": "",
          "post_date": "07/03/2024 13:28:10",
          "content": "<p>Thanks ! But whoosh runs indefinitely when I add this … Do you know why ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2902845,
              "author_name": "apparition",
              "author_url": "",
              "post_date": "07/03/2024 13:38:44",
              "content": "<p>Using wildcards (* and ?) makes queries run longer. I'm pretty sure that if every query contains a wildcard, then the 2500 queries will take more than the allowed hour. </p>\n<p>When I tested this specific subquery \"NOT cpc:*\", the query containing it would take about 15 seconds to run.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2913873,
      "author_name": "dhruvsandhu1",
      "author_url": "",
      "post_date": "07/09/2024 16:50:57",
      "content": "<p>Excellent work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2920931,
      "author_name": "dongwooim",
      "author_url": "",
      "post_date": "07/13/2024 23:14:56",
      "content": "<p>Nice work. Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2898783": "While researching whoosh, I discovered some features that can improve the LB score. Since rankings based on this knowledge seem to deviate from the competition's purpose, I’m sharing these insights.\n\n# 1. Tips for saving tokens\n## 1.1. Omit AND token\nhttps://whoosh.readthedocs.io/en/latest/querylang.html#boolean-operators\n- `ti:dog ti:cat` is equivalent to `ti:dog AND ti:cat`\n- `ti:dog ab:cat clm:fox` is equivalent to `ti:dog AND ab:cat AND clm:fox`\n\n## 1.2. Omit ADJ token\nhttps://whoosh.readthedocs.io/en/latest/querylang.html#individual-terms-and-phrases\n- `ti:\"open sesame\"` is equivalent to `ti:open ADJ1 ti:sesame`\n- `ti:\"open sesame\"~2` is equivalent to `ti:open ADJ2 ti:sesame`\n\n## 1.3. Search across fields\nhttps://whoosh.readthedocs.io/en/latest/parsing.html#letting-the-user-search-multiple-fields-by-default\n- `apple` is equivalent to `ti:apple OR ab:apple OR clm:apple OR detd:apple OR cpc:apple`\n\n# 2. Tips for result order\n- Results are returned in descending order of their TF-IDF scores\n- If the scores are the same, the results are returned in the order they were indexed in whoosh. Therefore, changing the order of the query does not affect the order of the results.\n\n# Examples\n```py\nimport whoosh_utils\n\ndocs = [\n    {\"publication_number\":\"one\", \"title\":\"dog dummy cat\", \"abstract\":\"sesame dummy open\", \"claims\":\"cat\", \"description\":\"dummy\", \"cpc\": [\"dummy\"]},\n    {\"publication_number\":\"two\", \"title\":\"cat dummy dummy dog\", \"abstract\":\"open sesame\", \"claims\":\"cat cat\", \"description\":\"cat\", \"cpc\": [\"dummy\"]},\n    {\"publication_number\":\"three\", \"title\":\"cat\", \"abstract\":\"open dummy dummy sesame\", \"claims\":\"cat cat cat\", \"description\":\"dog\", \"cpc\": [\"dummy\"]},\n]\nwhoosh_utils.create_index(f\"sample\", docs)\n\nsample_index = whoosh_utils.load_index(f\"sample\")\nsample_searcher = whoosh_utils.get_searcher(sample_index)\nqp = whoosh_utils.get_query_parser()\nqv = whoosh_utils.QueryValidator()\n\nfor query in [\"ti:dog ti:cat\", 'ab:\"open sesame\"', 'ab:\"open sesame\"~3', 'dog', 'clm:cat', 'detd:cat OR detd:dog', 'detd:dog OR detd:cat']:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    results = whoosh_utils.execute_query(query, qp, sample_searcher)\n    print(query, results, num_tokens)\n```\n\n```py\n# Example of checking \"1.1. Omit AND token\"\nti:dog ti:cat ['one', 'two'] 2\n# Example of checking \"1.2. Omit ADJ token\"\nab:\"open sesame\" ['two'] 2\nab:\"open sesame\"~3 ['two', 'three'] 2\n# Example of checking \"1.3. Search across fields\"\ndog ['three', 'one', 'two'] 1\n# Example of checking \"2. Tips for result order\"\nclm:cat ['three', 'two', 'one'] 1\ndetd:cat OR detd:dog ['two', 'three'] 3\ndetd:dog OR detd:cat ['two', 'three'] 3\n```\n\nHopefully this will help you to focus on the main part of the competition.",
    "2898799": "Thank you, valuable information!\nAre operators also counted before tokens are counted? (AND, OR, etc.)?",
    "2898823": "Thanks! How can we retrieve TF-IDF scores of all documents from the index?",
    "2898891": "https://www.kaggle.com/competitions/uspto-explainable-ai/overview\n> Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.\n\nIs the question asking if operators are counted as well? If so, yes, operators are a type of token. (In my understanding)",
    "2898897": "You can directly look at the contents of whoosh.searching.Hit\n\n```py\ndef execute_query(query: str, qp, searcher, results_limit=50) -> list:\n    if len(query) > 10_000:\n        raise ValueError('Query length at exceeds 10,000 characters.')\n    if 'id:' in query:\n        raise ValueError('Searching for specific patent IDs is banned.')\n\n    to_search = qp.parse(query)\n    results = searcher.search(to_search, limit=results_limit)\n    # results = [x['id'] for x in results]\n    results = [(x.score, x['id']) for x in results]\n    assert len(results) <= results_limit\n    return results\n\n...\n\nfor query in [\"ti:dog ti:cat\", 'ab:\"open sesame\"', 'ab:\"open sesame\"~3', 'dog', 'clm:cat', 'detd:cat OR detd:dog', 'detd:dog OR detd:cat']:\n    qv.validate_query(query)\n    num_tokens = whoosh_utils.count_query_tokens(query)\n    # results = whoosh_utils.execute_query(query, qp, sample_searcher)\n    results = execute_query(query, qp, sample_searcher)\n    print(query, results, num_tokens)\n```\n\n```py\nti:dog ti:cat [(1.7123179275482192, 'one'), (1.7123179275482192, 'two')] 2\nab:\"open sesame\" [(1.4246358550964382, 'two')] 2\nab:\"open sesame\"~3 [(1.4246358550964382, 'two'), (1.4246358550964382, 'three')] 2\ndog [(1.4054651081081644, 'three'), (1.0, 'one'), (1.0, 'two')] 1\nclm:cat [(2.136953782644657, 'three'), (1.4246358550964382, 'two'), (0.7123179275482191, 'one')] 1\ndetd:cat OR detd:dog [(1.4054651081081644, 'two'), (1.4054651081081644, 'three')] 3\ndetd:dog OR detd:cat [(1.4054651081081644, 'two'), (1.4054651081081644, 'three')] 3\n```\n\nFor more details: https://whoosh.readthedocs.io/en/latest/api/searching.html#whoosh.searching.Hit",
    "2899071": "Here is another one:\nNOT cpc:*  - this query returns patents with no CPC codes linked to them",
    "2902831": "Thanks ! But whoosh runs indefinitely when I add this ... Do you know why ?",
    "2902845": "Using wildcards (* and ?) makes queries run longer. I'm pretty sure that if every query contains a wildcard, then the 2500 queries will take more than the allowed hour. \n\nWhen I tested this specific subquery \"NOT cpc:*\", the query containing it would take about 15 seconds to run.",
    "2913873": "Excellent work",
    "2920931": "Nice work. Thanks"
  },
  "source": "meta"
}