{
  "id": 522200,
  "title": "4th Place Solution",
  "url": "/competitions/uspto-explainable-ai/writeups/penguin46-4th-place-solution",
  "author_name": "",
  "post_date": "2024-07-28T03:48:16.250Z",
  "votes": 44,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, congratulations to the winning teams.</p>\n<p>It is very unfortunate that there were some magics regarding query construction. I only became aware of them 10 days ago and have since spent a lot of time considering how to use them. </p>\n<p>I will explain the solution using magic (LB 0.98) and the solution not using magic (LB 0.91).</p>\n<h2>Solution with magic (LB 0.98)</h2>\n<p>You can save AND tokens by constructing the query as follows. The number of tokens in the example below is “1”.</p>\n<blockquote>\n  <p>ti:”token1”detd:”token2””token3”ti:”token4”’</p>\n</blockquote>\n<p>This is equivalent to:</p>\n<blockquote>\n  <p>ti:token1 AND detd:token2 AND token3 AND ti:token4</p>\n</blockquote>\n<p>In my solution, I divided the target patents into 25 pairs and connected up to 25 common tokens by AND tokens in order of decreasing number of occurrences. The final query has the following form:</p>\n<blockquote>\n  <p>(”token1”ti:”token2”…detd:”token25”) OR (”token26”ti:”token27”…detd:”token50”) OR …</p>\n</blockquote>\n<p>To determine which pairs to match, I used the product of the token occurrence probabilities. After matching, each partial query is run against the test-index, which contains all patents in test.csv, and if other than the two targets are hit, they are not used for the final query. Next, hit each unused target one by one with the extra tokens. (As in the case of pairs, I only greedily combine tokens that occur infrequently.) Finally, I added common tokens whenever possible to ensure that the length of the query did not exceed the upper limit of 10,000.</p>\n<h2>Solution without magic (LB 0.91)</h2>\n<p>Simulated annealing for queries in the following format:</p>\n<blockquote>\n  <p>cpc:CPC1(token1 OR token2 OR …) OR cpc:CPC2(…) … OR (token1 OR … OR tokenN)</p>\n</blockquote>\n<p>The evaluation function in SA is the expected value of mAP. This is not an exact expectation, but it works fast enough and accurately enough. Please check <code>state.py</code> to be released later for details. (<a href=\"https://github.com/penguin-prg/uspto-4th-place-solution/blob/main/src/solver/state.py#L189\" target=\"_blank\">here</a>)</p>\n<p>Pre-processing of this solution requires about several weeks in my environment. Various speedup and disk space-related tips were made, for examples:</p>\n<ul>\n<li>Pre-file tokenized text data</li>\n<li>Compress the file as .bz2</li>\n<li>Use a Key Value Store called leveldb<ul>\n<li>Directly handling .bz2 is faster and easier to parallelize, but uploading to the Kaggle Dataset is very cumbersome when the number of files is too large.</li></ul></li>\n<li>When using leveldb, the data itself is consolidated into a single text file, and the DB stores the range to be read<ul>\n<li>i.e, key → (leveldb) → range → (.txt.bz2) → data</li>\n<li>Storing the data directly in the DB has deteriorated the DB construction and access speed</li></ul></li>\n<li>Use hash to split the DB<ul>\n<li>When uploading to kaggle notebook, if the DB is larger than 20GB, split it into several DBs.</li>\n<li>In this case, a hash function is used to determine which DB a key belongs to. This eliminates the need for additional dict or DB.</li></ul></li>\n<li>For large pre-processing, save a flag to avoid repeating the completed process<ul>\n<li>Processes may be interrupted by errors due to OOM or unexpected corner cases.</li>\n<li>When resumed, processes that are flagged as completed will be skipped.</li></ul></li>\n<li>etc…</li>\n</ul>\n<h1>Conclusion</h1>\n<p>Again, it is very sad that such magics were uncovered at the end of the competition, as this was a very interesting task. I found several other magics in addition to this one. (e.g., a way to express an exact match of multiple tokens in a single token) As far as I know, there is no precedent of LB being recalculated after deadline in past competitions due to leaks or other defects, so LB will probably be finalized as is. Fortunately, however, as far as I can see from LB, most of the top participants did not use these magics until at least 2 weeks before the competition deadline. I look forward to reading their solutions without magics.</p>\n<h1>Code</h1>\n<ul>\n<li>final submission (0.98) : <a href=\"https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-magic-lb0-98\" target=\"_blank\">https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-magic-lb0-98</a></li>\n<li>final submission (0.91) : <a href=\"https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-o-magic-lb0-91\" target=\"_blank\">https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-o-magic-lb0-91</a></li>\n<li>github: <a href=\"https://github.com/penguin-prg/uspto-4th-place-solution\" target=\"_blank\">https://github.com/penguin-prg/uspto-4th-place-solution</a></li>\n</ul>",
  "messages": [
    {
      "id": "2935056",
      "postDate": "07/25/2024 00:09:06",
      "content": "<p>First of all, congratulations to the winning teams.</p>\n<p>It is very unfortunate that there were some magics regarding query construction. I only became aware of them 10 days ago and have since spent a lot of time considering how to use them. </p>\n<p>I will explain the solution using magic (LB 0.98) and the solution not using magic (LB 0.91).</p>\n<h2>Solution with magic (LB 0.98)</h2>\n<p>You can save AND tokens by constructing the query as follows. The number of tokens in the example below is “1”.</p>\n<blockquote>\n  <p>ti:”token1”detd:”token2””token3”ti:”token4”’</p>\n</blockquote>\n<p>This is equivalent to:</p>\n<blockquote>\n  <p>ti:token1 AND detd:token2 AND token3 AND ti:token4</p>\n</blockquote>\n<p>In my solution, I divided the target patents into 25 pairs and connected up to 25 common tokens by AND tokens in order of decreasing number of occurrences. The final query has the following form:</p>\n<blockquote>\n  <p>(”token1”ti:”token2”…detd:”token25”) OR (”token26”ti:”token27”…detd:”token50”) OR …</p>\n</blockquote>\n<p>To determine which pairs to match, I used the product of the token occurrence probabilities. After matching, each partial query is run against the test-index, which contains all patents in test.csv, and if other than the two targets are hit, they are not used for the final query. Next, hit each unused target one by one with the extra tokens. (As in the case of pairs, I only greedily combine tokens that occur infrequently.) Finally, I added common tokens whenever possible to ensure that the length of the query did not exceed the upper limit of 10,000.</p>\n<h2>Solution without magic (LB 0.91)</h2>\n<p>Simulated annealing for queries in the following format:</p>\n<blockquote>\n  <p>cpc:CPC1(token1 OR token2 OR …) OR cpc:CPC2(…) … OR (token1 OR … OR tokenN)</p>\n</blockquote>\n<p>The evaluation function in SA is the expected value of mAP. This is not an exact expectation, but it works fast enough and accurately enough. Please check <code>state.py</code> to be released later for details. (<a href=\"https://github.com/penguin-prg/uspto-4th-place-solution/blob/main/src/solver/state.py#L189\" target=\"_blank\">here</a>)</p>\n<p>Pre-processing of this solution requires about several weeks in my environment. Various speedup and disk space-related tips were made, for examples:</p>\n<ul>\n<li>Pre-file tokenized text data</li>\n<li>Compress the file as .bz2</li>\n<li>Use a Key Value Store called leveldb<ul>\n<li>Directly handling .bz2 is faster and easier to parallelize, but uploading to the Kaggle Dataset is very cumbersome when the number of files is too large.</li></ul></li>\n<li>When using leveldb, the data itself is consolidated into a single text file, and the DB stores the range to be read<ul>\n<li>i.e, key → (leveldb) → range → (.txt.bz2) → data</li>\n<li>Storing the data directly in the DB has deteriorated the DB construction and access speed</li></ul></li>\n<li>Use hash to split the DB<ul>\n<li>When uploading to kaggle notebook, if the DB is larger than 20GB, split it into several DBs.</li>\n<li>In this case, a hash function is used to determine which DB a key belongs to. This eliminates the need for additional dict or DB.</li></ul></li>\n<li>For large pre-processing, save a flag to avoid repeating the completed process<ul>\n<li>Processes may be interrupted by errors due to OOM or unexpected corner cases.</li>\n<li>When resumed, processes that are flagged as completed will be skipped.</li></ul></li>\n<li>etc…</li>\n</ul>\n<h1>Conclusion</h1>\n<p>Again, it is very sad that such magics were uncovered at the end of the competition, as this was a very interesting task. I found several other magics in addition to this one. (e.g., a way to express an exact match of multiple tokens in a single token) As far as I know, there is no precedent of LB being recalculated after deadline in past competitions due to leaks or other defects, so LB will probably be finalized as is. Fortunately, however, as far as I can see from LB, most of the top participants did not use these magics until at least 2 weeks before the competition deadline. I look forward to reading their solutions without magics.</p>\n<h1>Code</h1>\n<ul>\n<li>final submission (0.98) : <a href=\"https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-magic-lb0-98\" target=\"_blank\">https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-magic-lb0-98</a></li>\n<li>final submission (0.91) : <a href=\"https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-o-magic-lb0-91\" target=\"_blank\">https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-o-magic-lb0-91</a></li>\n<li>github: <a href=\"https://github.com/penguin-prg/uspto-4th-place-solution\" target=\"_blank\">https://github.com/penguin-prg/uspto-4th-place-solution</a></li>\n</ul>",
      "rawMarkdown": "First of all, congratulations to the winning teams.\n\nIt is very unfortunate that there were some magics regarding query construction. I only became aware of them 10 days ago and have since spent a lot of time considering how to use them. \n\nI will explain the solution using magic (LB 0.98) and the solution not using magic (LB 0.91).\n\n## Solution with magic (LB 0.98)\n\nYou can save AND tokens by constructing the query as follows. The number of tokens in the example below is “1”.\n\n> ti:”token1”detd:”token2””token3”ti:”token4”’\n\nThis is equivalent to:\n\n> ti:token1 AND detd:token2 AND token3 AND ti:token4\n\nIn my solution, I divided the target patents into 25 pairs and connected up to 25 common tokens by AND tokens in order of decreasing number of occurrences. The final query has the following form:\n\n> (”token1”ti:”token2”…detd:”token25”) OR (”token26”ti:”token27”…detd:”token50”) OR …\n\nTo determine which pairs to match, I used the product of the token occurrence probabilities. After matching, each partial query is run against the test-index, which contains all patents in test.csv, and if other than the two targets are hit, they are not used for the final query. Next, hit each unused target one by one with the extra tokens. (As in the case of pairs, I only greedily combine tokens that occur infrequently.) Finally, I added common tokens whenever possible to ensure that the length of the query did not exceed the upper limit of 10,000.\n\n## Solution without magic (LB 0.91)\n\nSimulated annealing for queries in the following format:\n\n> cpc:CPC1(token1 OR token2 OR …) OR cpc:CPC2(…) … OR (token1 OR … OR tokenN)\n> \n\nThe evaluation function in SA is the expected value of mAP. This is not an exact expectation, but it works fast enough and accurately enough. Please check `state.py` to be released later for details. ([here](https://github.com/penguin-prg/uspto-4th-place-solution/blob/main/src/solver/state.py#L189))\n\nPre-processing of this solution requires about several weeks in my environment. Various speedup and disk space-related tips were made, for examples:\n\n- Pre-file tokenized text data\n- Compress the file as .bz2\n- Use a Key Value Store called leveldb\n    - Directly handling .bz2 is faster and easier to parallelize, but uploading to the Kaggle Dataset is very cumbersome when the number of files is too large.\n- When using leveldb, the data itself is consolidated into a single text file, and the DB stores the range to be read\n    - i.e, key → (leveldb) → range → (.txt.bz2) → data\n    - Storing the data directly in the DB has deteriorated the DB construction and access speed\n- Use hash to split the DB\n    - When uploading to kaggle notebook, if the DB is larger than 20GB, split it into several DBs.\n    - In this case, a hash function is used to determine which DB a key belongs to. This eliminates the need for additional dict or DB.\n- For large pre-processing, save a flag to avoid repeating the completed process\n    - Processes may be interrupted by errors due to OOM or unexpected corner cases.\n    - When resumed, processes that are flagged as completed will be skipped.\n- etc…\n\n# Conclusion\n\nAgain, it is very sad that such magics were uncovered at the end of the competition, as this was a very interesting task. I found several other magics in addition to this one. (e.g., a way to express an exact match of multiple tokens in a single token) As far as I know, there is no precedent of LB being recalculated after deadline in past competitions due to leaks or other defects, so LB will probably be finalized as is. Fortunately, however, as far as I can see from LB, most of the top participants did not use these magics until at least 2 weeks before the competition deadline. I look forward to reading their solutions without magics.\n\n# Code\n\n- final submission (0.98) : https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-magic-lb0-98\n- final submission (0.91) : https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-o-magic-lb0-91\n- github: https://github.com/penguin-prg/uspto-4th-place-solution",
      "votes": null
    },
    {
      "id": "2935202",
      "postDate": "07/25/2024 04:04:06",
      "content": "<p>Congratulations on the 4th place!<br>\nI have a question about the index:</p>\n<blockquote>\n  <p>After matching, each partial query is run against the test-index, which contains all patents in test.csv</p>\n</blockquote>\n<p>Does it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?</p>",
      "rawMarkdown": "Congratulations on the 4th place!\nI have a question about the index:\n\n> After matching, each partial query is run against the test-index, which contains all patents in test.csv\n\nDoes it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?",
      "votes": null
    },
    {
      "id": "2935216",
      "postDate": "07/25/2024 04:41:05",
      "content": "<p>Thank you! And congratulations on your solo gold!</p>\n<blockquote>\n  <p>Does it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?</p>\n</blockquote>\n<p>No, it contains 125K patents. (2500 rows x 50 neighbors) </p>",
      "rawMarkdown": "Thank you! And congratulations on your solo gold!\n\n> Does it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?\n\nNo, it contains 125K patents. (2500 rows x 50 neighbors)",
      "votes": null
    },
    {
      "id": "2935372",
      "postDate": "07/25/2024 07:57:18",
      "content": "<p>Congrats.</p>\n<blockquote>\n  <p>there were some magics regarding query construction. I only became aware of them 10 days</p>\n</blockquote>\n<p>How did you become aware?</p>\n<p>Asking because I doubt I would have found it if I had entered the competition.</p>",
      "rawMarkdown": "Congrats.\n\n> there were some magics regarding query construction. I only became aware of them 10 days\n\nHow did you become aware?\n\nAsking because I doubt I would have found it if I had entered the competition.",
      "votes": null
    },
    {
      "id": "2935419",
      "postDate": "07/25/2024 08:28:29",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nThank you!</p>\n<blockquote>\n  <p>How did you become aware?</p>\n</blockquote>\n<p>The big trigger was the appearance of a participant who got 0.99 in LB. I was 0.91 at the time and couldn't believe that I could get 0.99 in a straightforward way, so I started looking for magic. Then I realized that the token count was calculated using the length of the list, with the query divided by spaces and parentheses, so I looked for other ways to make whoosh recognize the token boundaries and finally found the answer.</p>",
      "rawMarkdown": "cpmpml \nThank you!\n\n> How did you become aware?\n\nThe big trigger was the appearance of a participant who got 0.99 in LB. I was 0.91 at the time and couldn't believe that I could get 0.99 in a straightforward way, so I started looking for magic. Then I realized that the token count was calculated using the length of the list, with the query divided by spaces and parentheses, so I looked for other ways to make whoosh recognize the token boundaries and finally found the answer.",
      "votes": null
    },
    {
      "id": "2935674",
      "postDate": "07/25/2024 13:46:15",
      "content": "<p>Thats a great achievement ! Congratulations !</p>",
      "rawMarkdown": "Thats a great achievement ! Congratulations !",
      "votes": null
    },
    {
      "id": "2935781",
      "postDate": "07/25/2024 14:49:50",
      "content": "<p>Thanks. I am impressed by your score without magic.</p>",
      "rawMarkdown": "Thanks. I am impressed by your score without magic.",
      "votes": null
    },
    {
      "id": "2937257",
      "postDate": "07/26/2024 22:06:51",
      "content": "<p>Your score without magic is also impressive. Congratulations.</p>",
      "rawMarkdown": "Your score without magic is also impressive. Congratulations.",
      "votes": null
    },
    {
      "id": "2940849",
      "postDate": "07/30/2024 14:00:08",
      "content": "<p>Congratulations on your 4th place finish, <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a>!<br>\nYour approach using \"magic\" for optimizing AND tokens is impressive and demonstrates deep technical insight. The non-\"magic\" solution also highlights your ability to tackle complex problems creatively. Thanks for sharing your detailed solutions and preprocessing tips—they're valuable contributions to the community.</p>",
      "rawMarkdown": "Congratulations on your 4th place finish, @ryotayoshinobu!\nYour approach using \"magic\" for optimizing AND tokens is impressive and demonstrates deep technical insight. The non-\"magic\" solution also highlights your ability to tackle complex problems creatively. Thanks for sharing your detailed solutions and preprocessing tips—they're valuable contributions to the community.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2935202,
      "author_name": "apparition",
      "author_url": "",
      "post_date": "07/25/2024 04:04:06",
      "content": "<p>Congratulations on the 4th place!<br>\nI have a question about the index:</p>\n<blockquote>\n  <p>After matching, each partial query is run against the test-index, which contains all patents in test.csv</p>\n</blockquote>\n<p>Does it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2935216,
          "author_name": "ryotayoshinobu",
          "author_url": "",
          "post_date": "07/25/2024 04:41:05",
          "content": "<p>Thank you! And congratulations on your solo gold!</p>\n<blockquote>\n  <p>Does it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?</p>\n</blockquote>\n<p>No, it contains 125K patents. (2500 rows x 50 neighbors) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2935372,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/25/2024 07:57:18",
      "content": "<p>Congrats.</p>\n<blockquote>\n  <p>there were some magics regarding query construction. I only became aware of them 10 days</p>\n</blockquote>\n<p>How did you become aware?</p>\n<p>Asking because I doubt I would have found it if I had entered the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2935419,
          "author_name": "ryotayoshinobu",
          "author_url": "",
          "post_date": "07/25/2024 08:28:29",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nThank you!</p>\n<blockquote>\n  <p>How did you become aware?</p>\n</blockquote>\n<p>The big trigger was the appearance of a participant who got 0.99 in LB. I was 0.91 at the time and couldn't believe that I could get 0.99 in a straightforward way, so I started looking for magic. Then I realized that the token count was calculated using the length of the list, with the query divided by spaces and parentheses, so I looked for other ways to make whoosh recognize the token boundaries and finally found the answer.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2935781,
              "author_name": "cpmpml",
              "author_url": "",
              "post_date": "07/25/2024 14:49:50",
              "content": "<p>Thanks. I am impressed by your score without magic.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2935674,
      "author_name": "aadiar",
      "author_url": "",
      "post_date": "07/25/2024 13:46:15",
      "content": "<p>Thats a great achievement ! Congratulations !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2937257,
      "author_name": "asiful109",
      "author_url": "",
      "post_date": "07/26/2024 22:06:51",
      "content": "<p>Your score without magic is also impressive. Congratulations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2940849,
      "author_name": "",
      "author_url": "",
      "post_date": "07/30/2024 14:00:08",
      "content": "<p>Congratulations on your 4th place finish, <a href=\"https://www.kaggle.com/ryotayoshinobu\" target=\"_blank\">@ryotayoshinobu</a>!<br>\nYour approach using \"magic\" for optimizing AND tokens is impressive and demonstrates deep technical insight. The non-\"magic\" solution also highlights your ability to tackle complex problems creatively. Thanks for sharing your detailed solutions and preprocessing tips—they're valuable contributions to the community.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935056": "First of all, congratulations to the winning teams.\n\nIt is very unfortunate that there were some magics regarding query construction. I only became aware of them 10 days ago and have since spent a lot of time considering how to use them. \n\nI will explain the solution using magic (LB 0.98) and the solution not using magic (LB 0.91).\n\n## Solution with magic (LB 0.98)\n\nYou can save AND tokens by constructing the query as follows. The number of tokens in the example below is “1”.\n\n> ti:”token1”detd:”token2””token3”ti:”token4”’\n\nThis is equivalent to:\n\n> ti:token1 AND detd:token2 AND token3 AND ti:token4\n\nIn my solution, I divided the target patents into 25 pairs and connected up to 25 common tokens by AND tokens in order of decreasing number of occurrences. The final query has the following form:\n\n> (”token1”ti:”token2”…detd:”token25”) OR (”token26”ti:”token27”…detd:”token50”) OR …\n\nTo determine which pairs to match, I used the product of the token occurrence probabilities. After matching, each partial query is run against the test-index, which contains all patents in test.csv, and if other than the two targets are hit, they are not used for the final query. Next, hit each unused target one by one with the extra tokens. (As in the case of pairs, I only greedily combine tokens that occur infrequently.) Finally, I added common tokens whenever possible to ensure that the length of the query did not exceed the upper limit of 10,000.\n\n## Solution without magic (LB 0.91)\n\nSimulated annealing for queries in the following format:\n\n> cpc:CPC1(token1 OR token2 OR …) OR cpc:CPC2(…) … OR (token1 OR … OR tokenN)\n> \n\nThe evaluation function in SA is the expected value of mAP. This is not an exact expectation, but it works fast enough and accurately enough. Please check `state.py` to be released later for details. ([here](https://github.com/penguin-prg/uspto-4th-place-solution/blob/main/src/solver/state.py#L189))\n\nPre-processing of this solution requires about several weeks in my environment. Various speedup and disk space-related tips were made, for examples:\n\n- Pre-file tokenized text data\n- Compress the file as .bz2\n- Use a Key Value Store called leveldb\n    - Directly handling .bz2 is faster and easier to parallelize, but uploading to the Kaggle Dataset is very cumbersome when the number of files is too large.\n- When using leveldb, the data itself is consolidated into a single text file, and the DB stores the range to be read\n    - i.e, key → (leveldb) → range → (.txt.bz2) → data\n    - Storing the data directly in the DB has deteriorated the DB construction and access speed\n- Use hash to split the DB\n    - When uploading to kaggle notebook, if the DB is larger than 20GB, split it into several DBs.\n    - In this case, a hash function is used to determine which DB a key belongs to. This eliminates the need for additional dict or DB.\n- For large pre-processing, save a flag to avoid repeating the completed process\n    - Processes may be interrupted by errors due to OOM or unexpected corner cases.\n    - When resumed, processes that are flagged as completed will be skipped.\n- etc…\n\n# Conclusion\n\nAgain, it is very sad that such magics were uncovered at the end of the competition, as this was a very interesting task. I found several other magics in addition to this one. (e.g., a way to express an exact match of multiple tokens in a single token) As far as I know, there is no precedent of LB being recalculated after deadline in past competitions due to leaks or other defects, so LB will probably be finalized as is. Fortunately, however, as far as I can see from LB, most of the top participants did not use these magics until at least 2 weeks before the competition deadline. I look forward to reading their solutions without magics.\n\n# Code\n\n- final submission (0.98) : https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-magic-lb0-98\n- final submission (0.91) : https://www.kaggle.com/code/ryotayoshinobu/uspto-4th-place-solution-w-o-magic-lb0-91\n- github: https://github.com/penguin-prg/uspto-4th-place-solution",
    "2935202": "Congratulations on the 4th place!\nI have a question about the index:\n\n> After matching, each partial query is run against the test-index, which contains all patents in test.csv\n\nDoes it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?",
    "2935216": "Thank you! And congratulations on your solo gold!\n\n> Does it mean that you created a whoosh index containing all the 13M patents and used it in the submission script to test the subqueries?\n\nNo, it contains 125K patents. (2500 rows x 50 neighbors)",
    "2935372": "Congrats.\n\n> there were some magics regarding query construction. I only became aware of them 10 days\n\nHow did you become aware?\n\nAsking because I doubt I would have found it if I had entered the competition.",
    "2935419": "cpmpml \nThank you!\n\n> How did you become aware?\n\nThe big trigger was the appearance of a participant who got 0.99 in LB. I was 0.91 at the time and couldn't believe that I could get 0.99 in a straightforward way, so I started looking for magic. Then I realized that the token count was calculated using the length of the list, with the query divided by spaces and parentheses, so I looked for other ways to make whoosh recognize the token boundaries and finally found the answer.",
    "2935674": "Thats a great achievement ! Congratulations !",
    "2935781": "Thanks. I am impressed by your score without magic.",
    "2937257": "Your score without magic is also impressive. Congratulations.",
    "2940849": "Congratulations on your 4th place finish, @ryotayoshinobu!\nYour approach using \"magic\" for optimizing AND tokens is impressive and demonstrates deep technical insight. The non-\"magic\" solution also highlights your ability to tackle complex problems creatively. Thanks for sharing your detailed solutions and preprocessing tips—they're valuable contributions to the community."
  },
  "source": "meta"
}