{
  "id": 513097,
  "title": "CV Strategy",
  "url": "/competitions/uspto-explainable-ai/discussion/513097",
  "author_name": "Vincent Schuler",
  "post_date": "2024-06-18T15:10:18.863000",
  "votes": 9,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I would like to build a reliable CV strategy to properly tackle this competition but at first glance, it requires to load a lot of publications into the whoosh utils to avoid optimistic CV scores.</p>\n<p>If I'm right, it should be done here : <code>whoosh_idx = whoosh_utils.load_index(path_to_index)</code></p>\n<p>But loading a large number of publications will be very long, so I wonder how you adress this problem… ? Am I missing something ?</p>\n<p>Thanks !</p>",
  "messages": [
    {
      "id": 2877794,
      "postDate": "2024-06-18T15:10:18.863Z",
      "content": "<p>Hello,</p>\n<p>I would like to build a reliable CV strategy to properly tackle this competition but at first glance, it requires to load a lot of publications into the whoosh utils to avoid optimistic CV scores.</p>\n<p>If I'm right, it should be done here : <code>whoosh_idx = whoosh_utils.load_index(path_to_index)</code></p>\n<p>But loading a large number of publications will be very long, so I wonder how you adress this problem… ? Am I missing something ?</p>\n<p>Thanks !</p>",
      "rawMarkdown": "Hello,\n\nI would like to build a reliable CV strategy to properly tackle this competition but at first glance, it requires to load a lot of publications into the whoosh utils to avoid optimistic CV scores.\n\nIf I'm right, it should be done here : `whoosh_idx = whoosh_utils.load_index(path_to_index)`\n\nBut loading a large number of publications will be very long, so I wonder how you adress this problem... ? Am I missing something ?\n\nThanks !",
      "votes": 8
    },
    {
      "id": 2878906,
      "postDate": "2024-06-19T08:33:00.950Z",
      "content": "<p>I am calculating local scores on the following 200k patents.</p>\n<ul>\n<li>2500 patents since 1975.</li>\n<li>Their neighborhoods (~125k)</li>\n<li>Other random patents (about 75k)</li>\n</ul>",
      "rawMarkdown": "I am calculating local scores on the following 200k patents.\n- 2500 patents since 1975.\n- Their neighborhoods (~125k)\n- Other random patents (about 75k)\n\n",
      "votes": 6,
      "replies": [
        {
          "id": 2878910,
          "postDate": "2024-06-19T08:36:22.140Z",
          "content": "<p>But this still leaves a gap of about 0.05 between local scores and LB; there may be some bias in the sampling of test.csv and test index.</p>",
          "rawMarkdown": "But this still leaves a gap of about 0.05 between local scores and LB; there may be some bias in the sampling of test.csv and test index.",
          "votes": 2,
          "replies": [
            {
              "id": 2890504,
              "postDate": "2024-06-26T06:38:28.903Z",
              "content": "<p>just wondering what implementation of map50 you are using? trying to get a local validation score that is close to LB</p>",
              "rawMarkdown": "just wondering what implementation of map50 you are using? trying to get a local validation score that is close to LB"
            },
            {
              "id": 2890571,
              "postDate": "2024-06-26T07:10:41.720Z",
              "content": "<p><a href=\"https://www.kaggle.com/thedilpreet\" target=\"_blank\">@thedilpreet</a> I'm using this code.</p>\n<pre><code> () -&gt; :\n     (target_ids) == \n     (results) &lt; :\n        results.append()\n    hit = \n    ap = \n     i, result  (results):\n         result  target_ids:\n            hit += \n        ap += hit / (i + )\n    ap /= (target_ids)\n     ap\n</code></pre>",
              "rawMarkdown": "@thedilpreet I'm using this code.\n\n```python\ndef compute_ap(results: List[str], target_ids: List[str]) -> float:\n    assert len(target_ids) == 50\n    while len(results) < 50:\n        results.append(\"DUMMY\")\n    hit = 0\n    ap = 0\n    for i, result in enumerate(results):\n        if result in target_ids:\n            hit += 1\n        ap += hit / (i + 1)\n    ap /= len(target_ids)\n    return ap\n\n```",
              "votes": 5
            }
          ]
        },
        {
          "id": 2878918,
          "postDate": "2024-06-19T08:42:53.500Z",
          "content": "<p>Ok I see, thank you ! I imagine computing such a CV must be quite long, but at least it's representative.</p>",
          "rawMarkdown": "Ok I see, thank you ! I imagine computing such a CV must be quite long, but at least it's representative.",
          "votes": 2,
          "replies": [
            {
              "id": 2878921,
              "postDate": "2024-06-19T08:45:09.627Z",
              "content": "<p>In my environment the evaluation is fast enough for me. (2500 samples in less than 1 minute)</p>",
              "rawMarkdown": "In my environment the evaluation is fast enough for me. (2500 samples in less than 1 minute)",
              "votes": 2
            },
            {
              "id": 2879644,
              "postDate": "2024-06-19T17:53:18.603Z",
              "content": "<p>I'm doing something similar to penguin46, although I'm not including 75k random patents. If you want to optimize for speed I can advice you to build your own searcher optimized for the types of queries you want to run, I'm able to run queries in less than 0.2ms on average that way (=2500 queries in less than 500ms on a single thread).</p>",
              "rawMarkdown": "I'm doing something similar to penguin46, although I'm not including 75k random patents. If you want to optimize for speed I can advice you to build your own searcher optimized for the types of queries you want to run, I'm able to run queries in less than 0.2ms on average that way (=2500 queries in less than 500ms on a single thread).",
              "votes": 3
            },
            {
              "id": 2880665,
              "postDate": "2024-06-20T10:10:42.293Z",
              "content": "<p>I see, thanks to both of you !</p>",
              "rawMarkdown": "I see, thanks to both of you !",
              "votes": 2
            },
            {
              "id": 2908240,
              "postDate": "2024-07-06T07:12:00.187Z",
              "content": "<p>Can you share how to optimize searcher? I use<br>\n<code>searcher = whoosh_utils.get_searcher(whoosh_utils.load_index(\". /test_index\"))</code><br>\nbut found it very slow</p>",
              "rawMarkdown": "Can you share how to optimize searcher? I use\n`searcher = whoosh_utils.get_searcher(whoosh_utils.load_index(\". /test_index\")) `\nbut found it very slow"
            },
            {
              "id": 2908301,
              "postDate": "2024-07-06T08:25:22.917Z",
              "content": "<p>I am using the exact same code as you.</p>\n<pre><code>     = whoosh_utils.load_index(TRAIN_INDEX_PATH)\n     = whoosh_utils.get_searcher(train_idx)\n     = whoosh_utils.get_query_parser()\n</code></pre>",
              "rawMarkdown": "I am using the exact same code as you.\n```\n    train_idx = whoosh_utils.load_index(TRAIN_INDEX_PATH)\n    searcher = whoosh_utils.get_searcher(train_idx)\n    qp = whoosh_utils.get_query_parser()\n```",
              "votes": 1
            },
            {
              "id": 2908898,
              "postDate": "2024-07-06T16:09:50.980Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2908905,
              "postDate": "2024-07-06T16:16:41.627Z",
              "content": "<p>So do you have any way to speed up the evaluation process? It takes me more then 2h to execute 2,500 queries using this searcher</p>",
              "rawMarkdown": "So do you have any way to speed up the evaluation process? It takes me more then 2h to execute 2,500 queries using this searcher"
            },
            {
              "id": 2908922,
              "postDate": "2024-07-06T16:29:23.853Z",
              "content": "<p>No, there is no special tricks. In my experiments, <code>execute_query</code> finishes within 30sec locally and within 2min in the kaggle notebook.</p>",
              "rawMarkdown": "No, there is no special tricks. In my experiments, `execute_query` finishes within 30sec locally and within 2min in the kaggle notebook.",
              "votes": 1
            },
            {
              "id": 2908930,
              "postDate": "2024-07-06T16:35:13.143Z",
              "content": "<p>Okay, thank you for your response. I guess my queries are too complex. Do you mind sharing how to build the optimized searcher? <a href=\"https://www.kaggle.com/jmerle\" target=\"_blank\">@jmerle</a> </p>",
              "rawMarkdown": "Okay, thank you for your response. I guess my queries are too complex. Do you mind sharing how to build the optimized searcher? @jmerle ",
              "votes": 1
            },
            {
              "id": 2909237,
              "postDate": "2024-07-06T19:32:59.093Z",
              "content": "<p>What makes my searcher fast are bitsets, optimized data storage, and (probably) C++. I won't share my code during the competition, but I can share a summary.</p>\n<p>I assign each patent in my index an id starting from 0, and store my index data in such a way that I can efficiently retrieve the bitsets of each possible term. If bit X is set in a term's bitset, the patent with id X matches the term. Executing a query then comes down to reading bitsets and applying bitwise operations between them.</p>\n<p>If there are &lt;=50 set bits left after these operations we are done. If there are &gt;50 set bits left we have to sort the matching patents using the TF-IDF weighting scheme. The most costly part of this is retrieving the number of times each term matches each patent, although the impact of that depends a lot on where and how your index is stored.</p>\n<p>My &lt;0.2ms per query times were on an in-memory index containing 200k patents. In-memory indexes are ideal, but not always feasible depending on how much patents you want your index to contain and what query constructs you want to support.</p>\n<p>To limit the size of the index and to optimize performance my searcher is currently limited to OR, AND, NOT, and XOR operators. It also only accepts terms of the format <code>&lt;category&gt;:&lt;token&gt;</code>. where the category may be one of <code>cpc</code>, <code>ti</code>, <code>ab</code>, <code>clm</code>, and <code>detd</code>, and where the token may not contain spaces. It doesn't support proximity operators, wildcards, multi-word terms, or uncategorized terms.</p>",
              "rawMarkdown": "What makes my searcher fast are bitsets, optimized data storage, and (probably) C++. I won't share my code during the competition, but I can share a summary.\n\nI assign each patent in my index an id starting from 0, and store my index data in such a way that I can efficiently retrieve the bitsets of each possible term. If bit X is set in a term's bitset, the patent with id X matches the term. Executing a query then comes down to reading bitsets and applying bitwise operations between them.\n\nIf there are <=50 set bits left after these operations we are done. If there are >50 set bits left we have to sort the matching patents using the TF-IDF weighting scheme. The most costly part of this is retrieving the number of times each term matches each patent, although the impact of that depends a lot on where and how your index is stored.\n\nMy <0.2ms per query times were on an in-memory index containing 200k patents. In-memory indexes are ideal, but not always feasible depending on how much patents you want your index to contain and what query constructs you want to support.\n\nTo limit the size of the index and to optimize performance my searcher is currently limited to OR, AND, NOT, and XOR operators. It also only accepts terms of the format `<category>:<token>`. where the category may be one of `cpc`, `ti`, `ab`, `clm`, and `detd`, and where the token may not contain spaces. It doesn't support proximity operators, wildcards, multi-word terms, or uncategorized terms.",
              "votes": 1
            },
            {
              "id": 2914520,
              "postDate": "2024-07-10T02:18:40.667Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2914523,
              "postDate": "2024-07-10T02:20:00.533Z",
              "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> May I ask you, \"within 30sec locally and within 2min in the kaggle notebook.\" refers to the test set of text.csv, which is the time of 10 samples? Because I found that my “save and run all” took about 3 minutes, and it might take 5-6 hours to complete the score, I felt that was too long.😢</p>",
              "rawMarkdown": "@ryotayoshinobu May I ask you, \"within 30sec locally and within 2min in the kaggle notebook.\" refers to the test set of text.csv, which is the time of 10 samples? Because I found that my “save and run all” took about 3 minutes, and it might take 5-6 hours to complete the score, I felt that was too long.😢"
            },
            {
              "id": 2916358,
              "postDate": "2024-07-11T00:00:46.513Z",
              "content": "<p>\"30sec and 2min\" are time to run <code>execute_query</code> for 2500 queries. They do not include any other time.</p>",
              "rawMarkdown": "\"30sec and 2min\" are time to run `execute_query` for 2500 queries. They do not include any other time.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2880021,
          "postDate": "2024-06-20T02:05:06.647Z",
          "content": "<p>Very interesting. I am eager for competition end to learn how you have done this. I have not even a clue on how you could do anything more than optimized keywords with or logic for this comp similar to the public nb</p>",
          "rawMarkdown": "Very interesting. I am eager for competition end to learn how you have done this. I have not even a clue on how you could do anything more than optimized keywords with or logic for this comp similar to the public nb"
        },
        {
          "id": 2903652,
          "postDate": "2024-07-03T23:22:03.660Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> Are you taking a random sample of 2500 patents after 1975?</p>",
          "rawMarkdown": "@ryotayoshinobu Are you taking a random sample of 2500 patents after 1975?",
          "replies": [
            {
              "id": 2903653,
              "postDate": "2024-07-03T23:24:21.907Z",
              "content": "<p><a href=\"https://www.kaggle.com/chiragmahapatra\" target=\"_blank\">@chiragmahapatra</a> <br>\nyes.</p>",
              "rawMarkdown": "@chiragmahapatra \nyes."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2878906,
      "author_name": "penguin46",
      "author_url": "",
      "post_date": "2024-06-19T08:33:00.950000",
      "content": "<p>I am calculating local scores on the following 200k patents.</p>\n<ul>\n<li>2500 patents since 1975.</li>\n<li>Their neighborhoods (~125k)</li>\n<li>Other random patents (about 75k)</li>\n</ul>",
      "votes": 6,
      "replies": [
        {
          "id": 2878910,
          "author_name": "penguin46",
          "author_url": "",
          "post_date": "2024-06-19T08:36:22.140000",
          "content": "<p>But this still leaves a gap of about 0.05 between local scores and LB; there may be some bias in the sampling of test.csv and test index.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2890504,
              "author_name": "Dilpreet",
              "author_url": "",
              "post_date": "2024-06-26T06:38:28.903000",
              "content": "<p>just wondering what implementation of map50 you are using? trying to get a local validation score that is close to LB</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2890571,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-06-26T07:10:41.720000",
              "content": "<p><a href=\"https://www.kaggle.com/thedilpreet\" target=\"_blank\">@thedilpreet</a> I'm using this code.</p>\n<pre><code> () -&gt; :\n     (target_ids) == \n     (results) &lt; :\n        results.append()\n    hit = \n    ap = \n     i, result  (results):\n         result  target_ids:\n            hit += \n        ap += hit / (i + )\n    ap /= (target_ids)\n     ap\n</code></pre>",
              "votes": 5,
              "replies": []
            }
          ]
        },
        {
          "id": 2878918,
          "author_name": "Vincent Schuler",
          "author_url": "",
          "post_date": "2024-06-19T08:42:53.500000",
          "content": "<p>Ok I see, thank you ! I imagine computing such a CV must be quite long, but at least it's representative.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2878921,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-06-19T08:45:09.627000",
              "content": "<p>In my environment the evaluation is fast enough for me. (2500 samples in less than 1 minute)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2879644,
              "author_name": "Jasper",
              "author_url": "",
              "post_date": "2024-06-19T17:53:18.603000",
              "content": "<p>I'm doing something similar to penguin46, although I'm not including 75k random patents. If you want to optimize for speed I can advice you to build your own searcher optimized for the types of queries you want to run, I'm able to run queries in less than 0.2ms on average that way (=2500 queries in less than 500ms on a single thread).</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2880665,
              "author_name": "Vincent Schuler",
              "author_url": "",
              "post_date": "2024-06-20T10:10:42.293000",
              "content": "<p>I see, thanks to both of you !</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2908240,
              "author_name": "Filtered",
              "author_url": "",
              "post_date": "2024-07-06T07:12:00.187000",
              "content": "<p>Can you share how to optimize searcher? I use<br>\n<code>searcher = whoosh_utils.get_searcher(whoosh_utils.load_index(\". /test_index\"))</code><br>\nbut found it very slow</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2908301,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-07-06T08:25:22.917000",
              "content": "<p>I am using the exact same code as you.</p>\n<pre><code>     = whoosh_utils.load_index(TRAIN_INDEX_PATH)\n     = whoosh_utils.get_searcher(train_idx)\n     = whoosh_utils.get_query_parser()\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2908898,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-07-06T16:09:50.980000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2908905,
              "author_name": "Filtered",
              "author_url": "",
              "post_date": "2024-07-06T16:16:41.627000",
              "content": "<p>So do you have any way to speed up the evaluation process? It takes me more then 2h to execute 2,500 queries using this searcher</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2908922,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-07-06T16:29:23.853000",
              "content": "<p>No, there is no special tricks. In my experiments, <code>execute_query</code> finishes within 30sec locally and within 2min in the kaggle notebook.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2908930,
              "author_name": "Filtered",
              "author_url": "",
              "post_date": "2024-07-06T16:35:13.143000",
              "content": "<p>Okay, thank you for your response. I guess my queries are too complex. Do you mind sharing how to build the optimized searcher? <a href=\"https://www.kaggle.com/jmerle\" target=\"_blank\">@jmerle</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2909237,
              "author_name": "Jasper",
              "author_url": "",
              "post_date": "2024-07-06T19:32:59.093000",
              "content": "<p>What makes my searcher fast are bitsets, optimized data storage, and (probably) C++. I won't share my code during the competition, but I can share a summary.</p>\n<p>I assign each patent in my index an id starting from 0, and store my index data in such a way that I can efficiently retrieve the bitsets of each possible term. If bit X is set in a term's bitset, the patent with id X matches the term. Executing a query then comes down to reading bitsets and applying bitwise operations between them.</p>\n<p>If there are &lt;=50 set bits left after these operations we are done. If there are &gt;50 set bits left we have to sort the matching patents using the TF-IDF weighting scheme. The most costly part of this is retrieving the number of times each term matches each patent, although the impact of that depends a lot on where and how your index is stored.</p>\n<p>My &lt;0.2ms per query times were on an in-memory index containing 200k patents. In-memory indexes are ideal, but not always feasible depending on how much patents you want your index to contain and what query constructs you want to support.</p>\n<p>To limit the size of the index and to optimize performance my searcher is currently limited to OR, AND, NOT, and XOR operators. It also only accepts terms of the format <code>&lt;category&gt;:&lt;token&gt;</code>. where the category may be one of <code>cpc</code>, <code>ti</code>, <code>ab</code>, <code>clm</code>, and <code>detd</code>, and where the token may not contain spaces. It doesn't support proximity operators, wildcards, multi-word terms, or uncategorized terms.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2914520,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-07-10T02:18:40.667000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2914523,
              "author_name": "Alannikos",
              "author_url": "",
              "post_date": "2024-07-10T02:20:00.533000",
              "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> May I ask you, \"within 30sec locally and within 2min in the kaggle notebook.\" refers to the test set of text.csv, which is the time of 10 samples? Because I found that my “save and run all” took about 3 minutes, and it might take 5-6 hours to complete the score, I felt that was too long.😢</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2916358,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-07-11T00:00:46.513000",
              "content": "<p>\"30sec and 2min\" are time to run <code>execute_query</code> for 2500 queries. They do not include any other time.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2880021,
          "author_name": "Fir_las_47",
          "author_url": "",
          "post_date": "2024-06-20T02:05:06.647000",
          "content": "<p>Very interesting. I am eager for competition end to learn how you have done this. I have not even a clue on how you could do anything more than optimized keywords with or logic for this comp similar to the public nb</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2903652,
          "author_name": "ChiragMahapatra",
          "author_url": "",
          "post_date": "2024-07-03T23:22:03.660000",
          "content": "<p><a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a> Are you taking a random sample of 2500 patents after 1975?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2903653,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-07-03T23:24:21.907000",
              "content": "<p><a href=\"https://www.kaggle.com/chiragmahapatra\" target=\"_blank\">@chiragmahapatra</a> <br>\nyes.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2877794": "Hello,\n\nI would like to build a reliable CV strategy to properly tackle this competition but at first glance, it requires to load a lot of publications into the whoosh utils to avoid optimistic CV scores.\n\nIf I'm right, it should be done here : `whoosh_idx = whoosh_utils.load_index(path_to_index)`\n\nBut loading a large number of publications will be very long, so I wonder how you adress this problem... ? Am I missing something ?\n\nThanks !",
    "2878906": "I am calculating local scores on the following 200k patents.\n- 2500 patents since 1975.\n- Their neighborhoods (~125k)\n- Other random patents (about 75k)\n\n"
  }
}