{
  "id": 522303,
  "title": "Fixing submission scoring error due to long query time",
  "url": "/competitions/uspto-explainable-ai/discussion/522303",
  "author_name": "dt",
  "post_date": "2024-07-25T13:15:36.467000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>2 weeks before the deadline I started having submission scoring errors again after updating my dataset. After extensive checks I was sure that it was due to query time, and I rushed to fix this without reverting my data. Here’s my findings and I hope it will help all those who face this issue to finally resolve the problem or at least get submissions with not too degraded results.</p>\n<h6>Some background about my queries</h6>\n<p>My queries are all disjunctions of bigrams/trigrams from title, abstract, claims and description fields (the approach is detailed <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522301\" target=\"_blank\">here</a>). Parsing is fast and there are no illegal characters (I strip special characters in preprocessing).<br>\nIn Whoosh, I believe all phrase terms are actually converted to adjacent queries under the hood. Majority of the slowness I think comes from querying phrases with very high frequency individual words in very large fields (description!), as there are a very large number of positions to iterate through. Not all phrases with high frequency words are slow though (in fact the majority are not). <br>\nWhoosh doesn’t have a query profiler like Elasticsearch, and I didn’t investigate the actual query execution so I don’t have an answer on how to check whether a query is slow or not without running it. Phrases in title, abstract, claims are safe regardless of word frequency. I think Whoosh also has some caching for example a field cache that can distort the raw query time of a term (in general all search engines have multiple caching layers). Overall I think if you’re using expensive operators (positional, wildcard) on dense fields you may need to drop them.</p>\n<p>I found that my slow phrases are consistently slow across different sampled indexes of size 2500, which gave me confidence on dropping bad phrases just based on single query execution. If I had a lot of time and resources I could test the query time of each phrase, but in the end I settled for only checking phrases that only consist of the top 250 words from my phrase vocabulary. This was sufficient to cut about a second off my mean query time, a second off the standard deviation, and bring 97%-99% of my test queries to below 8s.</p>\n<h6>Competition info</h6>\n<p>The following information have been mentioned in the competition overview:</p>\n<blockquote>\n  <p>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel.  </p>\n</blockquote>\n<p>The environment is unknown and it’s hard to evaluate the performance of Whoosh parallel search, but I believe the limitations are closer to 60 mins for all sequential rather than 4x60. I’m also not sure if it's a shared environment but I suspect so as my running times were slightly longer near to the competition deadline.</p>\n<p>In practice I found these were some guideline numbers that were ‘safe’ and ‘not safe’ for me, for various validation indexes with 2500 targets:<br>\nsafe: mean query time of 2.1s, std dev of 1.4s, max of 10s, total time of 87 mins<br>\nunsafe: mean query time of 3.5s, std dev of 2.7s, max of 20s, total time of 145 mins</p>\n<h6>Handling slow queries without failing submission completely</h6>\n<p>In whoosh_utils execute_query() we can implement a searcher with a TimeLimitCollector. This forces the searcher to raise an exception and return once the given time limit is reached. </p>\n<pre><code> () -&gt; :\n     (query) &gt; :\n         ValueError()\n       query:\n         ValueError()\n\n    tc = TopCollector(limit=results_limit)\n    tec = TermsCollector(tc)\n    tlc = TimeLimitCollector(tec, timelimit=time_limit)\n\n    to_search = qp.parse(query)\n    results = (, )\n    :\n        searcher.search_with_collector(to_search, tlc)\n        results = ([x[]  x  tlc.results()], [k  k  tec.termdocs.keys()])\n         results\n     TimeLimit:\n        results = ([x[]  x  tlc.results()], [k  k  tec.termdocs.keys()])\n         TimeLimit(, results)\n     Exception  e:\n         e\n</code></pre>\n<p>Note the TermsCollector. This enables us to get partial results even if time limit is reached. With these partial results we can cross reference the original query to find the ‘safe’ words.</p>\n<p>Note however the partial results are given as unigrams, which makes it hard to correlate exactly which <em>phrases</em> are safe. Here I just give a naive approach of checking if all the phrase words are contained within the partial results. Also, if the validation index does not contain that word, it gives a false negative. But at least in this way, we can have some partial query instead of dropping the query completely when time limit is reached (or using some random heuristic)</p>\n<p>With the combination of dropping individual phrases with bad query time and implementing time limit with partial results, I was able to climb back slowly from the scoring error black hole to a stable state.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fe3fb7fd0773b64a9427524c8c016abc4%2FScreenshot%202024-07-25%20at%209.11.39PM.png?generation=1721913150365151&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F3a50c6e77cd7e85f3ee5b8bfbe99c027%2FScreenshot%202024-07-25%20at%209.12.01PM.png?generation=1721913166903252&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2935640,
      "postDate": "2024-07-25T13:15:36.467Z",
      "content": "<p>2 weeks before the deadline I started having submission scoring errors again after updating my dataset. After extensive checks I was sure that it was due to query time, and I rushed to fix this without reverting my data. Here’s my findings and I hope it will help all those who face this issue to finally resolve the problem or at least get submissions with not too degraded results.</p>\n<h6>Some background about my queries</h6>\n<p>My queries are all disjunctions of bigrams/trigrams from title, abstract, claims and description fields (the approach is detailed <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522301\" target=\"_blank\">here</a>). Parsing is fast and there are no illegal characters (I strip special characters in preprocessing).<br>\nIn Whoosh, I believe all phrase terms are actually converted to adjacent queries under the hood. Majority of the slowness I think comes from querying phrases with very high frequency individual words in very large fields (description!), as there are a very large number of positions to iterate through. Not all phrases with high frequency words are slow though (in fact the majority are not). <br>\nWhoosh doesn’t have a query profiler like Elasticsearch, and I didn’t investigate the actual query execution so I don’t have an answer on how to check whether a query is slow or not without running it. Phrases in title, abstract, claims are safe regardless of word frequency. I think Whoosh also has some caching for example a field cache that can distort the raw query time of a term (in general all search engines have multiple caching layers). Overall I think if you’re using expensive operators (positional, wildcard) on dense fields you may need to drop them.</p>\n<p>I found that my slow phrases are consistently slow across different sampled indexes of size 2500, which gave me confidence on dropping bad phrases just based on single query execution. If I had a lot of time and resources I could test the query time of each phrase, but in the end I settled for only checking phrases that only consist of the top 250 words from my phrase vocabulary. This was sufficient to cut about a second off my mean query time, a second off the standard deviation, and bring 97%-99% of my test queries to below 8s.</p>\n<h6>Competition info</h6>\n<p>The following information have been mentioned in the competition overview:</p>\n<blockquote>\n  <p>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel.  </p>\n</blockquote>\n<p>The environment is unknown and it’s hard to evaluate the performance of Whoosh parallel search, but I believe the limitations are closer to 60 mins for all sequential rather than 4x60. I’m also not sure if it's a shared environment but I suspect so as my running times were slightly longer near to the competition deadline.</p>\n<p>In practice I found these were some guideline numbers that were ‘safe’ and ‘not safe’ for me, for various validation indexes with 2500 targets:<br>\nsafe: mean query time of 2.1s, std dev of 1.4s, max of 10s, total time of 87 mins<br>\nunsafe: mean query time of 3.5s, std dev of 2.7s, max of 20s, total time of 145 mins</p>\n<h6>Handling slow queries without failing submission completely</h6>\n<p>In whoosh_utils execute_query() we can implement a searcher with a TimeLimitCollector. This forces the searcher to raise an exception and return once the given time limit is reached. </p>\n<pre><code> () -&gt; :\n     (query) &gt; :\n         ValueError()\n       query:\n         ValueError()\n\n    tc = TopCollector(limit=results_limit)\n    tec = TermsCollector(tc)\n    tlc = TimeLimitCollector(tec, timelimit=time_limit)\n\n    to_search = qp.parse(query)\n    results = (, )\n    :\n        searcher.search_with_collector(to_search, tlc)\n        results = ([x[]  x  tlc.results()], [k  k  tec.termdocs.keys()])\n         results\n     TimeLimit:\n        results = ([x[]  x  tlc.results()], [k  k  tec.termdocs.keys()])\n         TimeLimit(, results)\n     Exception  e:\n         e\n</code></pre>\n<p>Note the TermsCollector. This enables us to get partial results even if time limit is reached. With these partial results we can cross reference the original query to find the ‘safe’ words.</p>\n<p>Note however the partial results are given as unigrams, which makes it hard to correlate exactly which <em>phrases</em> are safe. Here I just give a naive approach of checking if all the phrase words are contained within the partial results. Also, if the validation index does not contain that word, it gives a false negative. But at least in this way, we can have some partial query instead of dropping the query completely when time limit is reached (or using some random heuristic)</p>\n<p>With the combination of dropping individual phrases with bad query time and implementing time limit with partial results, I was able to climb back slowly from the scoring error black hole to a stable state.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fe3fb7fd0773b64a9427524c8c016abc4%2FScreenshot%202024-07-25%20at%209.11.39PM.png?generation=1721913150365151&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F3a50c6e77cd7e85f3ee5b8bfbe99c027%2FScreenshot%202024-07-25%20at%209.12.01PM.png?generation=1721913166903252&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "2 weeks before the deadline I started having submission scoring errors again after updating my dataset. After extensive checks I was sure that it was due to query time, and I rushed to fix this without reverting my data. Here’s my findings and I hope it will help all those who face this issue to finally resolve the problem or at least get submissions with not too degraded results.\n\n###### Some background about my queries\nMy queries are all disjunctions of bigrams/trigrams from title, abstract, claims and description fields (the approach is detailed [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522301)). Parsing is fast and there are no illegal characters (I strip special characters in preprocessing).\nIn Whoosh, I believe all phrase terms are actually converted to adjacent queries under the hood. Majority of the slowness I think comes from querying phrases with very high frequency individual words in very large fields (description!), as there are a very large number of positions to iterate through. Not all phrases with high frequency words are slow though (in fact the majority are not). \nWhoosh doesn’t have a query profiler like Elasticsearch, and I didn’t investigate the actual query execution so I don’t have an answer on how to check whether a query is slow or not without running it. Phrases in title, abstract, claims are safe regardless of word frequency. I think Whoosh also has some caching for example a field cache that can distort the raw query time of a term (in general all search engines have multiple caching layers). Overall I think if you’re using expensive operators (positional, wildcard) on dense fields you may need to drop them.\n\nI found that my slow phrases are consistently slow across different sampled indexes of size 2500, which gave me confidence on dropping bad phrases just based on single query execution. If I had a lot of time and resources I could test the query time of each phrase, but in the end I settled for only checking phrases that only consist of the top 250 words from my phrase vocabulary. This was sufficient to cut about a second off my mean query time, a second off the standard deviation, and bring 97%-99% of my test queries to below 8s.\n\n###### Competition info\nThe following information have been mentioned in the competition overview:\n>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel.  \n\nThe environment is unknown and it’s hard to evaluate the performance of Whoosh parallel search, but I believe the limitations are closer to 60 mins for all sequential rather than 4x60. I’m also not sure if it's a shared environment but I suspect so as my running times were slightly longer near to the competition deadline.\n\nIn practice I found these were some guideline numbers that were ‘safe’ and ‘not safe’ for me, for various validation indexes with 2500 targets:\nsafe: mean query time of 2.1s, std dev of 1.4s, max of 10s, total time of 87 mins\nunsafe: mean query time of 3.5s, std dev of 2.7s, max of 20s, total time of 145 mins\n\n###### Handling slow queries without failing submission completely\nIn whoosh_utils execute_query() we can implement a searcher with a TimeLimitCollector. This forces the searcher to raise an exception and return once the given time limit is reached. \n```python\ndef execute_query(query: str, qp, searcher, time_limit, results_limit=50) -> list:\n    if len(query) > 10_000:\n        raise ValueError('Query length at exceeds 10,000 characters.')\n    if 'id:' in query:\n        raise ValueError('Searching for specific patent IDs is banned.')\n    \n    tc = TopCollector(limit=results_limit)\n    tec = TermsCollector(tc)\n    tlc = TimeLimitCollector(tec, timelimit=time_limit)\n    \n    to_search = qp.parse(query)\n    results = (None, None)\n    try:\n        searcher.search_with_collector(to_search, tlc)\n        results = ([x['id'] for x in tlc.results()], [k for k in tec.termdocs.keys()])\n        return results\n    except TimeLimit:\n        results = ([x['id'] for x in tlc.results()], [k for k in tec.termdocs.keys()])\n        raise TimeLimit(f'Query exceeds time limit of {time_limit}s', results)\n    except Exception as e:\n        raise e\n```\nNote the TermsCollector. This enables us to get partial results even if time limit is reached. With these partial results we can cross reference the original query to find the ‘safe’ words.\n\nNote however the partial results are given as unigrams, which makes it hard to correlate exactly which _phrases_ are safe. Here I just give a naive approach of checking if all the phrase words are contained within the partial results. Also, if the validation index does not contain that word, it gives a false negative. But at least in this way, we can have some partial query instead of dropping the query completely when time limit is reached (or using some random heuristic)\n\nWith the combination of dropping individual phrases with bad query time and implementing time limit with partial results, I was able to climb back slowly from the scoring error black hole to a stable state.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fe3fb7fd0773b64a9427524c8c016abc4%2FScreenshot%202024-07-25%20at%209.11.39PM.png?generation=1721913150365151&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F3a50c6e77cd7e85f3ee5b8bfbe99c027%2FScreenshot%202024-07-25%20at%209.12.01PM.png?generation=1721913166903252&alt=media)",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2935640": "2 weeks before the deadline I started having submission scoring errors again after updating my dataset. After extensive checks I was sure that it was due to query time, and I rushed to fix this without reverting my data. Here’s my findings and I hope it will help all those who face this issue to finally resolve the problem or at least get submissions with not too degraded results.\n\n###### Some background about my queries\nMy queries are all disjunctions of bigrams/trigrams from title, abstract, claims and description fields (the approach is detailed [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/522301)). Parsing is fast and there are no illegal characters (I strip special characters in preprocessing).\nIn Whoosh, I believe all phrase terms are actually converted to adjacent queries under the hood. Majority of the slowness I think comes from querying phrases with very high frequency individual words in very large fields (description!), as there are a very large number of positions to iterate through. Not all phrases with high frequency words are slow though (in fact the majority are not). \nWhoosh doesn’t have a query profiler like Elasticsearch, and I didn’t investigate the actual query execution so I don’t have an answer on how to check whether a query is slow or not without running it. Phrases in title, abstract, claims are safe regardless of word frequency. I think Whoosh also has some caching for example a field cache that can distort the raw query time of a term (in general all search engines have multiple caching layers). Overall I think if you’re using expensive operators (positional, wildcard) on dense fields you may need to drop them.\n\nI found that my slow phrases are consistently slow across different sampled indexes of size 2500, which gave me confidence on dropping bad phrases just based on single query execution. If I had a lot of time and resources I could test the query time of each phrase, but in the end I settled for only checking phrases that only consist of the top 250 words from my phrase vocabulary. This was sufficient to cut about a second off my mean query time, a second off the standard deviation, and bring 97%-99% of my test queries to below 8s.\n\n###### Competition info\nThe following information have been mentioned in the competition overview:\n>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel.  \n\nThe environment is unknown and it’s hard to evaluate the performance of Whoosh parallel search, but I believe the limitations are closer to 60 mins for all sequential rather than 4x60. I’m also not sure if it's a shared environment but I suspect so as my running times were slightly longer near to the competition deadline.\n\nIn practice I found these were some guideline numbers that were ‘safe’ and ‘not safe’ for me, for various validation indexes with 2500 targets:\nsafe: mean query time of 2.1s, std dev of 1.4s, max of 10s, total time of 87 mins\nunsafe: mean query time of 3.5s, std dev of 2.7s, max of 20s, total time of 145 mins\n\n###### Handling slow queries without failing submission completely\nIn whoosh_utils execute_query() we can implement a searcher with a TimeLimitCollector. This forces the searcher to raise an exception and return once the given time limit is reached. \n```python\ndef execute_query(query: str, qp, searcher, time_limit, results_limit=50) -> list:\n    if len(query) > 10_000:\n        raise ValueError('Query length at exceeds 10,000 characters.')\n    if 'id:' in query:\n        raise ValueError('Searching for specific patent IDs is banned.')\n    \n    tc = TopCollector(limit=results_limit)\n    tec = TermsCollector(tc)\n    tlc = TimeLimitCollector(tec, timelimit=time_limit)\n    \n    to_search = qp.parse(query)\n    results = (None, None)\n    try:\n        searcher.search_with_collector(to_search, tlc)\n        results = ([x['id'] for x in tlc.results()], [k for k in tec.termdocs.keys()])\n        return results\n    except TimeLimit:\n        results = ([x['id'] for x in tlc.results()], [k for k in tec.termdocs.keys()])\n        raise TimeLimit(f'Query exceeds time limit of {time_limit}s', results)\n    except Exception as e:\n        raise e\n```\nNote the TermsCollector. This enables us to get partial results even if time limit is reached. With these partial results we can cross reference the original query to find the ‘safe’ words.\n\nNote however the partial results are given as unigrams, which makes it hard to correlate exactly which _phrases_ are safe. Here I just give a naive approach of checking if all the phrase words are contained within the partial results. Also, if the validation index does not contain that word, it gives a false negative. But at least in this way, we can have some partial query instead of dropping the query completely when time limit is reached (or using some random heuristic)\n\nWith the combination of dropping individual phrases with bad query time and implementing time limit with partial results, I was able to climb back slowly from the scoring error black hole to a stable state.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2Fe3fb7fd0773b64a9427524c8c016abc4%2FScreenshot%202024-07-25%20at%209.11.39PM.png?generation=1721913150365151&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4269349%2F3a50c6e77cd7e85f3ee5b8bfbe99c027%2FScreenshot%202024-07-25%20at%209.12.01PM.png?generation=1721913166903252&alt=media)"
  }
}