{
  "id": 522258,
  "title": "2nd place solution",
  "url": "/competitions/uspto-explainable-ai/writeups/shun-pi-2nd-place-solution",
  "author_name": "",
  "post_date": "2024-07-30T09:30:54.270Z",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to the host and congratulations to the winners.<br>\nI'm honestly very disappointed to have missed out on 1st place, but I'm glad to win my first prize and 5th gold medal on this competition!</p>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>solution with magic</td>\n<td>0.99934</td>\n<td>0.99849</td>\n<td>0.99872</td>\n</tr>\n<tr>\n<td>solution without magic</td>\n<td>0.89992</td>\n<td>0.90295</td>\n<td>0.90427</td>\n</tr>\n</tbody>\n</table>\n<h2>Magic</h2>\n<p>There are two (maybe unnoticed by the host) magics in this competition, and I used the former in my solution.</p>\n<ol>\n<li>no whitespace AND<br>\n<code>ti:\"abcd\"ab:\"efgh\"clm:\"ijkl\"</code></li>\n<li>no whitespace n-gram<br>\n<code>ti:\"abcd/efgh/ijkl\"</code> (there are several possible delimiters other than the slash)</li>\n</ol>\n<p>If you notice one of the two magics, you can get a score above 0.8 even with a simple solution like identifying 1 patent in 1 subquery.<br>\nWithout magic, the best submitted LB was 0.90. However, since I found magic, I focused on improving the solution using it, and I think there is much room for improvement in the solution without magic.</p>\n<h2>Insights on Metrics and Test Dataset</h2>\n<ul>\n<li>I suspect that the test_index contains non-targets that are similar to target intentionally. This can be guessed from the fact that when you submit a solution using a query that may include several non-targets, you get a large CV-LB gap, e.g., CV:0.8 vs. LB:0.5.</li>\n<li>According to the Whoosh Docs, the order of query results is ordered using TF-IDF scores, and it is difficult to control this order.</li>\n<li>This competition's metric has the nature that query results that do not include non-targets will have significantly higher scores. For example, the expected value of metric for 25 targets and 25 non-targets is 0.50 (if the order is random as said in previous discussion), but for 25 targets and 0 non-targets it is 0.84.</li>\n<li>From the above discussion, we can see that the solution to construct a query that always has zero non-targets is good.</li>\n<li>To achieve this, we can aim for a solution like \"construct an index with data from all patents, and make a query that is determined to contain no non-targets\". However, it is difficult to build Whoosh with all patent data, so it is necessary to build my own fast search algorithm.</li>\n</ul>\n<h2>Validation</h2>\n<ul>\n<li>I obtained a high correlation with LB by constructing the following CV<ul>\n<li>Randomly select 2,500 rows from 1975 or later where title/abstract/claims/description are non-null.</li>\n<li>Based on the above discussion, limit queries to those that do not include non-targets and absorb the CV/LB gap caused by intentional addition of non-targets in LB.</li></ul></li>\n</ul>\n<h2>Solution</h2>\n<ul>\n<li><p>In the final solution, the query looks like this</p>\n<ul>\n<li><code>query = (subquery1 OR subquery2 OR …) NOT (negative-subquery1 OR negative-subquery2 OR …)</code></li>\n<li>Here, subquery and negative-subquery is several words connected by AND.</li>\n<li>Negative-subquery is used to cancel out the non-targets that appear in subquery. Using negative-subquery improved the CV by about 0.0002.</li></ul></li>\n<li><p>The process is divided into three main parts: \"constructing index\", \"generating subquery candidates\", and \"selecting a subquery to use in a query\".</p>\n<ul>\n<li>The \"constructing index\" is done in advance, and the rest of the processing is done for each of the 2500 rows.</li>\n<li>For time and memory efficiency, all of the following solution methods were implemented in C++.<ul>\n<li>Compared to the Python solution, it is about 5 times faster in time and 3 times more efficient in memory.</li></ul></li></ul></li>\n<li><p>Below is an illustration of the entire process.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fe49dfd63ae109f96c6a25abebfcf68b2%2F2024-07-30%20182835.png?generation=1722331846448753&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong>Constructing index</strong></p>\n<ul>\n<li>Scan all patent sentences for cpc/title/abstract/claims/description and make a two-dimensional variable-length list from each patent to word, and from each word to each patent.</li>\n<li>Words and patents are not treated as strings, but are all converted to integer IDs in advance (for memory efficiency).</li>\n<li>Since the data for claims and descriptions is so large that it cannot be stored in RAM, I deleted words with high frequency of occurrence and left words with low frequency of occurrence, thus reducing the data volume to 13% for claims and 1.5% for descriptions.</li>\n<li>C++ programs will use about 25GB of RAM, which is enough to satisfy the Notebook's 30GB limit.</li></ul></li>\n<li><p><strong>Generating subquery candidates</strong></p>\n<ul>\n<li>Create subquery candidates corresponding to 50 single target and 50*49/2=1225 target pairs.</li>\n<li>Create a base word set as follows.<ul>\n<li>For a single target, the base word set is all the words it contains.</li>\n<li>For a target pair, the base word set is all the words common to both targets.</li></ul></li>\n<li>If a subquery is simply a set of base words connected by AND, the query will take a long time to execute.<ul>\n<li>Since the intersection set calculation takes min(len(A), len(B)).</li></ul></li>\n<li>Therefore, it is faster to sort in ascending order of the size of the patent corresponding to each word in advance, and to compute the intersection set in that order.</li>\n<li>Also, the computation of the  intersection set can be speeded up by terminating the computation as soon as non-targets are removed from the set. This also improves the score, since it is sometimes possible to retrieve targets larger than two.</li>\n<li>If the set size is too large at the first computation of the intersection set, give up.</li>\n<li>If non-targets remain after a certain number of words, check if it is possible to cancel non-targets using negative-subquery, and if that is not possible, give up.</li>\n<li>In this way, the number and order of words to construct a subquery can be optimized, and the results of executing that subquery can be retrieved at the same time.</li></ul></li>\n<li><p><strong>Selecting a subquery to use in a query</strong></p>\n<ul>\n<li>Repeat the following to greedily add subqueries until all 50 targets are covered or the 25 subquery limits are used up.<ul>\n<li>Score all subquery candidates and select the highest scoring subquery.</li>\n<li>The score function is as follows.<ul>\n<li>$$score(\\text{subquery}) = (\\text{Number of new targets that can be covered by selecting this subquery}) - \\sum\\limits_{\\text{targets covered by this subquery}} (\\text{Number of times the target appears in all subquery candidates}) * 10^{-9}$$</li></ul></li>\n<li>The second term is used for tie-breaking between targets with the same number of new targets that can be covered. The more difficult a target is to cover (fewer candidates), the higher the score when it is covered.</li></ul></li>\n<li>In the final submission, this greedy method was improved very slightly using beam search (CV+0.00003).</li></ul></li>\n<li><p>Histogram of the number of targets that could be retrieved in the final sub is as follows.</p>\n<ul>\n<li>91% of the data gave a perfect score (50 targets)</li>\n<li>99% of the data gave more than 45 targets<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fcbe617f06b8231ec355b94ea87d1e652%2F2024-07-25%20172310.png?generation=1721895803039500&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Query Example</h2>\n<ul>\n<li><p>solution with magic(50 targets)</p>\n<ul>\n<li><code>(ab:\"srm\"clm:\"srm\"cpc:\"G01N33/6848\"clm:\"quantified\") OR (detd:\"collisionally\"cpc:\"G01N33/6848\"detd:\"collisional\"ab:\"spectrometry\"detd:\"spectrom\"clm:\"spectrometry\"clm:\"tandem\"ab:\"proteins\") OR (detd:\"higgs\"ab:\"spectrometry\"clm:\"peptides\"clm:\"ms\"ab:\"proteins\"clm:\"proteins\"ab:\"protein\") OR (detd:\"xics\"detd:\"silac\"detd:\"itraq\"detd:\"xic\"detd:\"sequest\"ab:\"quantifying\") OR (detd:\"massbank\"detd:\"metlin\"detd:\"hmdb\"detd:\"synapt\") OR (detd:\"desolvated\"ab:\"peaks\"clm:\"spectrometry\"clm:\"spectra\") OR (clm:\"maldi\"detd:\"ftms\"clm:\"electrospray\"detd:\"quadrupoles\"ab:\"spectrometry\"detd:\"spectrom\"ab:\"mass\"ti:\"methods\") OR (detd:\"muddiman\"detd:\"gygi\"detd:\"lysc\") OR (detd:\"lumos\"cpc:\"G16B40/10\"detd:\"lysc\") OR (cpc:\"G16B20/00\"ti:\"complex\"ti:\"sample\"ab:\"ion\") OR (ab:\"spectrometry\"clm:\"spectrometry\"ab:\"peptide\"ab:\"mass\"ab:\"detected\"ab:\"improved\") OR (detd:\"picotip\"detd:\"fibrinopeptide\"clm:\"trna\") OR (detd:\"chait\"detd:\"sequest\"detd:\"proteomes\"detd:\"endoproteinase\") OR (detd:\"spectrom\"clm:\"spectrometry\"clm:\"spectra\"ab:\"peptide\"ab:\"mass\"ab:\"data\"ab:\"invention\"ab:\"one\"ab:\"with\") OR (cpc:\"Y10T436/24\"clm:\"labeled\"ab:\"sample\"ab:\"cell\"ab:\"obtained\") OR (detd:\"silac\"detd:\"itraq\"ti:\"multiplexed\"detd:\"iodoacetyl\") OR (clm:\"biomolecules\"clm:\"fragmenting\"clm:\"biomolecule\"clm:\"abundance\"clm:\"desorption\"clm:\"ionization\"clm:\"spectra\"clm:\"peptides\"clm:\"spectrometer\"clm:\"assisted\"clm:\"proteins\") OR (cpc:\"G01N33/6848\"ti:\"spectrometry\"ti:\"mass\"clm:\"peptides\"ab:\"peptide\"ab:\"mass\"ab:\"sample\"ab:\"determining\"ab:\"present\"ab:\"least\") OR (clm:\"fentomole\")</code></li></ul></li>\n<li><p>solution without magic(39 targets)</p>\n<ul>\n<li><code>(clm:\"walwffk\") OR (detd:\"gruhler\" detd:\"qstar\") OR (detd:\"hdmse\" detd:\"roepstorff\") OR (detd:\"gillet\" detd:\"electrosprayed\") OR (detd:\"sofeware\" detd:\"3.4.21.8\") OR (ab:\"srm\" clm:\"srm\" cpc:\"G01N33/6848\" clm:\"quantified\") OR (clm:\"lysophatidylinositol\") OR (detd:\"econometrics\" detd:\"silac\") OR (detd:\"sequest\" ab:\"spectrometry\" clm:\"multidimensional\") OR (detd:\"lumos\" cpc:\"G16B40/10\" detd:\"lysc\") OR (detd:\"phosphoproteomes\" detd:\"picotti\") OR (detd:\"pp4\" detd:\"femtomole\") OR (detd:\"c10h9n6o2\") OR (detd:\"mortz\" cpc:\"H01J49/00\") OR (detd:\"aqyneiqgwdhlsllp\") OR (clm:\"proteome\" detd:\"photodissociation\") OR (detd:\"albar\" detd:\"silac\")</code></li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "2935414",
      "postDate": "07/25/2024 08:24:34",
      "content": "<p>Thanks to the host and congratulations to the winners.<br>\nI'm honestly very disappointed to have missed out on 1st place, but I'm glad to win my first prize and 5th gold medal on this competition!</p>\n<h2>Summary</h2>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>solution with magic</td>\n<td>0.99934</td>\n<td>0.99849</td>\n<td>0.99872</td>\n</tr>\n<tr>\n<td>solution without magic</td>\n<td>0.89992</td>\n<td>0.90295</td>\n<td>0.90427</td>\n</tr>\n</tbody>\n</table>\n<h2>Magic</h2>\n<p>There are two (maybe unnoticed by the host) magics in this competition, and I used the former in my solution.</p>\n<ol>\n<li>no whitespace AND<br>\n<code>ti:\"abcd\"ab:\"efgh\"clm:\"ijkl\"</code></li>\n<li>no whitespace n-gram<br>\n<code>ti:\"abcd/efgh/ijkl\"</code> (there are several possible delimiters other than the slash)</li>\n</ol>\n<p>If you notice one of the two magics, you can get a score above 0.8 even with a simple solution like identifying 1 patent in 1 subquery.<br>\nWithout magic, the best submitted LB was 0.90. However, since I found magic, I focused on improving the solution using it, and I think there is much room for improvement in the solution without magic.</p>\n<h2>Insights on Metrics and Test Dataset</h2>\n<ul>\n<li>I suspect that the test_index contains non-targets that are similar to target intentionally. This can be guessed from the fact that when you submit a solution using a query that may include several non-targets, you get a large CV-LB gap, e.g., CV:0.8 vs. LB:0.5.</li>\n<li>According to the Whoosh Docs, the order of query results is ordered using TF-IDF scores, and it is difficult to control this order.</li>\n<li>This competition's metric has the nature that query results that do not include non-targets will have significantly higher scores. For example, the expected value of metric for 25 targets and 25 non-targets is 0.50 (if the order is random as said in previous discussion), but for 25 targets and 0 non-targets it is 0.84.</li>\n<li>From the above discussion, we can see that the solution to construct a query that always has zero non-targets is good.</li>\n<li>To achieve this, we can aim for a solution like \"construct an index with data from all patents, and make a query that is determined to contain no non-targets\". However, it is difficult to build Whoosh with all patent data, so it is necessary to build my own fast search algorithm.</li>\n</ul>\n<h2>Validation</h2>\n<ul>\n<li>I obtained a high correlation with LB by constructing the following CV<ul>\n<li>Randomly select 2,500 rows from 1975 or later where title/abstract/claims/description are non-null.</li>\n<li>Based on the above discussion, limit queries to those that do not include non-targets and absorb the CV/LB gap caused by intentional addition of non-targets in LB.</li></ul></li>\n</ul>\n<h2>Solution</h2>\n<ul>\n<li><p>In the final solution, the query looks like this</p>\n<ul>\n<li><code>query = (subquery1 OR subquery2 OR …) NOT (negative-subquery1 OR negative-subquery2 OR …)</code></li>\n<li>Here, subquery and negative-subquery is several words connected by AND.</li>\n<li>Negative-subquery is used to cancel out the non-targets that appear in subquery. Using negative-subquery improved the CV by about 0.0002.</li></ul></li>\n<li><p>The process is divided into three main parts: \"constructing index\", \"generating subquery candidates\", and \"selecting a subquery to use in a query\".</p>\n<ul>\n<li>The \"constructing index\" is done in advance, and the rest of the processing is done for each of the 2500 rows.</li>\n<li>For time and memory efficiency, all of the following solution methods were implemented in C++.<ul>\n<li>Compared to the Python solution, it is about 5 times faster in time and 3 times more efficient in memory.</li></ul></li></ul></li>\n<li><p>Below is an illustration of the entire process.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fe49dfd63ae109f96c6a25abebfcf68b2%2F2024-07-30%20182835.png?generation=1722331846448753&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong>Constructing index</strong></p>\n<ul>\n<li>Scan all patent sentences for cpc/title/abstract/claims/description and make a two-dimensional variable-length list from each patent to word, and from each word to each patent.</li>\n<li>Words and patents are not treated as strings, but are all converted to integer IDs in advance (for memory efficiency).</li>\n<li>Since the data for claims and descriptions is so large that it cannot be stored in RAM, I deleted words with high frequency of occurrence and left words with low frequency of occurrence, thus reducing the data volume to 13% for claims and 1.5% for descriptions.</li>\n<li>C++ programs will use about 25GB of RAM, which is enough to satisfy the Notebook's 30GB limit.</li></ul></li>\n<li><p><strong>Generating subquery candidates</strong></p>\n<ul>\n<li>Create subquery candidates corresponding to 50 single target and 50*49/2=1225 target pairs.</li>\n<li>Create a base word set as follows.<ul>\n<li>For a single target, the base word set is all the words it contains.</li>\n<li>For a target pair, the base word set is all the words common to both targets.</li></ul></li>\n<li>If a subquery is simply a set of base words connected by AND, the query will take a long time to execute.<ul>\n<li>Since the intersection set calculation takes min(len(A), len(B)).</li></ul></li>\n<li>Therefore, it is faster to sort in ascending order of the size of the patent corresponding to each word in advance, and to compute the intersection set in that order.</li>\n<li>Also, the computation of the  intersection set can be speeded up by terminating the computation as soon as non-targets are removed from the set. This also improves the score, since it is sometimes possible to retrieve targets larger than two.</li>\n<li>If the set size is too large at the first computation of the intersection set, give up.</li>\n<li>If non-targets remain after a certain number of words, check if it is possible to cancel non-targets using negative-subquery, and if that is not possible, give up.</li>\n<li>In this way, the number and order of words to construct a subquery can be optimized, and the results of executing that subquery can be retrieved at the same time.</li></ul></li>\n<li><p><strong>Selecting a subquery to use in a query</strong></p>\n<ul>\n<li>Repeat the following to greedily add subqueries until all 50 targets are covered or the 25 subquery limits are used up.<ul>\n<li>Score all subquery candidates and select the highest scoring subquery.</li>\n<li>The score function is as follows.<ul>\n<li>$$score(\\text{subquery}) = (\\text{Number of new targets that can be covered by selecting this subquery}) - \\sum\\limits_{\\text{targets covered by this subquery}} (\\text{Number of times the target appears in all subquery candidates}) * 10^{-9}$$</li></ul></li>\n<li>The second term is used for tie-breaking between targets with the same number of new targets that can be covered. The more difficult a target is to cover (fewer candidates), the higher the score when it is covered.</li></ul></li>\n<li>In the final submission, this greedy method was improved very slightly using beam search (CV+0.00003).</li></ul></li>\n<li><p>Histogram of the number of targets that could be retrieved in the final sub is as follows.</p>\n<ul>\n<li>91% of the data gave a perfect score (50 targets)</li>\n<li>99% of the data gave more than 45 targets<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fcbe617f06b8231ec355b94ea87d1e652%2F2024-07-25%20172310.png?generation=1721895803039500&amp;alt=media\" alt=\"\"></li></ul></li>\n</ul>\n<h2>Query Example</h2>\n<ul>\n<li><p>solution with magic(50 targets)</p>\n<ul>\n<li><code>(ab:\"srm\"clm:\"srm\"cpc:\"G01N33/6848\"clm:\"quantified\") OR (detd:\"collisionally\"cpc:\"G01N33/6848\"detd:\"collisional\"ab:\"spectrometry\"detd:\"spectrom\"clm:\"spectrometry\"clm:\"tandem\"ab:\"proteins\") OR (detd:\"higgs\"ab:\"spectrometry\"clm:\"peptides\"clm:\"ms\"ab:\"proteins\"clm:\"proteins\"ab:\"protein\") OR (detd:\"xics\"detd:\"silac\"detd:\"itraq\"detd:\"xic\"detd:\"sequest\"ab:\"quantifying\") OR (detd:\"massbank\"detd:\"metlin\"detd:\"hmdb\"detd:\"synapt\") OR (detd:\"desolvated\"ab:\"peaks\"clm:\"spectrometry\"clm:\"spectra\") OR (clm:\"maldi\"detd:\"ftms\"clm:\"electrospray\"detd:\"quadrupoles\"ab:\"spectrometry\"detd:\"spectrom\"ab:\"mass\"ti:\"methods\") OR (detd:\"muddiman\"detd:\"gygi\"detd:\"lysc\") OR (detd:\"lumos\"cpc:\"G16B40/10\"detd:\"lysc\") OR (cpc:\"G16B20/00\"ti:\"complex\"ti:\"sample\"ab:\"ion\") OR (ab:\"spectrometry\"clm:\"spectrometry\"ab:\"peptide\"ab:\"mass\"ab:\"detected\"ab:\"improved\") OR (detd:\"picotip\"detd:\"fibrinopeptide\"clm:\"trna\") OR (detd:\"chait\"detd:\"sequest\"detd:\"proteomes\"detd:\"endoproteinase\") OR (detd:\"spectrom\"clm:\"spectrometry\"clm:\"spectra\"ab:\"peptide\"ab:\"mass\"ab:\"data\"ab:\"invention\"ab:\"one\"ab:\"with\") OR (cpc:\"Y10T436/24\"clm:\"labeled\"ab:\"sample\"ab:\"cell\"ab:\"obtained\") OR (detd:\"silac\"detd:\"itraq\"ti:\"multiplexed\"detd:\"iodoacetyl\") OR (clm:\"biomolecules\"clm:\"fragmenting\"clm:\"biomolecule\"clm:\"abundance\"clm:\"desorption\"clm:\"ionization\"clm:\"spectra\"clm:\"peptides\"clm:\"spectrometer\"clm:\"assisted\"clm:\"proteins\") OR (cpc:\"G01N33/6848\"ti:\"spectrometry\"ti:\"mass\"clm:\"peptides\"ab:\"peptide\"ab:\"mass\"ab:\"sample\"ab:\"determining\"ab:\"present\"ab:\"least\") OR (clm:\"fentomole\")</code></li></ul></li>\n<li><p>solution without magic(39 targets)</p>\n<ul>\n<li><code>(clm:\"walwffk\") OR (detd:\"gruhler\" detd:\"qstar\") OR (detd:\"hdmse\" detd:\"roepstorff\") OR (detd:\"gillet\" detd:\"electrosprayed\") OR (detd:\"sofeware\" detd:\"3.4.21.8\") OR (ab:\"srm\" clm:\"srm\" cpc:\"G01N33/6848\" clm:\"quantified\") OR (clm:\"lysophatidylinositol\") OR (detd:\"econometrics\" detd:\"silac\") OR (detd:\"sequest\" ab:\"spectrometry\" clm:\"multidimensional\") OR (detd:\"lumos\" cpc:\"G16B40/10\" detd:\"lysc\") OR (detd:\"phosphoproteomes\" detd:\"picotti\") OR (detd:\"pp4\" detd:\"femtomole\") OR (detd:\"c10h9n6o2\") OR (detd:\"mortz\" cpc:\"H01J49/00\") OR (detd:\"aqyneiqgwdhlsllp\") OR (clm:\"proteome\" detd:\"photodissociation\") OR (detd:\"albar\" detd:\"silac\")</code></li></ul></li>\n</ul>",
      "rawMarkdown": "Thanks to the host and congratulations to the winners.\nI'm honestly very disappointed to have missed out on 1st place, but I'm glad to win my first prize and 5th gold medal on this competition!\n\n## Summary\n\n|  | CV | Public LB | Private LB |\n| --- | --- | --- | --- |\n| solution with magic | 0.99934 | 0.99849 | 0.99872 |\n| solution without magic | 0.89992 | 0.90295 | 0.90427 |\n\n## Magic\n\nThere are two (maybe unnoticed by the host) magics in this competition, and I used the former in my solution.\n\n1. no whitespace AND\n  `ti:\"abcd\"ab:\"efgh\"clm:\"ijkl\"`\n2. no whitespace n-gram\n  `ti:\"abcd/efgh/ijkl\"` (there are several possible delimiters other than the slash)\n\nIf you notice one of the two magics, you can get a score above 0.8 even with a simple solution like identifying 1 patent in 1 subquery.\nWithout magic, the best submitted LB was 0.90. However, since I found magic, I focused on improving the solution using it, and I think there is much room for improvement in the solution without magic.\n\n## Insights on Metrics and Test Dataset\n\n-  I suspect that the test_index contains non-targets that are similar to target intentionally. This can be guessed from the fact that when you submit a solution using a query that may include several non-targets, you get a large CV-LB gap, e.g., CV:0.8 vs. LB:0.5.\n-  According to the Whoosh Docs, the order of query results is ordered using TF-IDF scores, and it is difficult to control this order.\n-  This competition's metric has the nature that query results that do not include non-targets will have significantly higher scores. For example, the expected value of metric for 25 targets and 25 non-targets is 0.50 (if the order is random as said in previous discussion), but for 25 targets and 0 non-targets it is 0.84.\n-  From the above discussion, we can see that the solution to construct a query that always has zero non-targets is good.\n-  To achieve this, we can aim for a solution like \"construct an index with data from all patents, and make a query that is determined to contain no non-targets\". However, it is difficult to build Whoosh with all patent data, so it is necessary to build my own fast search algorithm.\n\n## Validation\n\n- I obtained a high correlation with LB by constructing the following CV\n    - Randomly select 2,500 rows from 1975 or later where title/abstract/claims/description are non-null.\n    - Based on the above discussion, limit queries to those that do not include non-targets and absorb the CV/LB gap caused by intentional addition of non-targets in LB.\n\n## Solution\n\n- In the final solution, the query looks like this\n    - `query = (subquery1 OR subquery2 OR …) NOT (negative-subquery1 OR negative-subquery2 OR …)`\n    - Here, subquery and negative-subquery is several words connected by AND.\n    - Negative-subquery is used to cancel out the non-targets that appear in subquery. Using negative-subquery improved the CV by about 0.0002.\n\n- The process is divided into three main parts: \"constructing index\", \"generating subquery candidates\", and \"selecting a subquery to use in a query\".\n    - The \"constructing index\" is done in advance, and the rest of the processing is done for each of the 2500 rows.\n    - For time and memory efficiency, all of the following solution methods were implemented in C++.\n        - Compared to the Python solution, it is about 5 times faster in time and 3 times more efficient in memory.\n\n- Below is an illustration of the entire process.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fe49dfd63ae109f96c6a25abebfcf68b2%2F2024-07-30%20182835.png?generation=1722331846448753&alt=media)\n\n- **Constructing index**\n    - Scan all patent sentences for cpc/title/abstract/claims/description and make a two-dimensional variable-length list from each patent to word, and from each word to each patent.\n    - Words and patents are not treated as strings, but are all converted to integer IDs in advance (for memory efficiency).\n    - Since the data for claims and descriptions is so large that it cannot be stored in RAM, I deleted words with high frequency of occurrence and left words with low frequency of occurrence, thus reducing the data volume to 13% for claims and 1.5% for descriptions.\n    - C++ programs will use about 25GB of RAM, which is enough to satisfy the Notebook's 30GB limit.\n\n- **Generating subquery candidates**\n    - Create subquery candidates corresponding to 50 single target and 50*49/2=1225 target pairs.\n    - Create a base word set as follows.\n         - For a single target, the base word set is all the words it contains.\n         - For a target pair, the base word set is all the words common to both targets.\n    - If a subquery is simply a set of base words connected by AND, the query will take a long time to execute.\n         - Since the intersection set calculation takes min(len(A), len(B)).\n    - Therefore, it is faster to sort in ascending order of the size of the patent corresponding to each word in advance, and to compute the intersection set in that order.\n    - Also, the computation of the  intersection set can be speeded up by terminating the computation as soon as non-targets are removed from the set. This also improves the score, since it is sometimes possible to retrieve targets larger than two.\n    - If the set size is too large at the first computation of the intersection set, give up.\n    - If non-targets remain after a certain number of words, check if it is possible to cancel non-targets using negative-subquery, and if that is not possible, give up.\n    - In this way, the number and order of words to construct a subquery can be optimized, and the results of executing that subquery can be retrieved at the same time.\n\n- **Selecting a subquery to use in a query**\n    - Repeat the following to greedily add subqueries until all 50 targets are covered or the 25 subquery limits are used up.\n        - Score all subquery candidates and select the highest scoring subquery.\n        - The score function is as follows.\n             - $$score(\\text{subquery}) = (\\text{Number of new targets that can be covered by selecting this subquery}) - \\sum\\limits_{\\text{targets covered by this subquery}} (\\text{Number of times the target appears in all subquery candidates}) * 10^{-9}$$\n        - The second term is used for tie-breaking between targets with the same number of new targets that can be covered. The more difficult a target is to cover (fewer candidates), the higher the score when it is covered.\n    - In the final submission, this greedy method was improved very slightly using beam search (CV+0.00003).\n\n- Histogram of the number of targets that could be retrieved in the final sub is as follows.\n    - 91% of the data gave a perfect score (50 targets)\n    - 99% of the data gave more than 45 targets\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fcbe617f06b8231ec355b94ea87d1e652%2F2024-07-25%20172310.png?generation=1721895803039500&alt=media)\n\n## Query Example\n\n- solution with magic(50 targets)\n  - `(ab:\"srm\"clm:\"srm\"cpc:\"G01N33/6848\"clm:\"quantified\") OR (detd:\"collisionally\"cpc:\"G01N33/6848\"detd:\"collisional\"ab:\"spectrometry\"detd:\"spectrom\"clm:\"spectrometry\"clm:\"tandem\"ab:\"proteins\") OR (detd:\"higgs\"ab:\"spectrometry\"clm:\"peptides\"clm:\"ms\"ab:\"proteins\"clm:\"proteins\"ab:\"protein\") OR (detd:\"xics\"detd:\"silac\"detd:\"itraq\"detd:\"xic\"detd:\"sequest\"ab:\"quantifying\") OR (detd:\"massbank\"detd:\"metlin\"detd:\"hmdb\"detd:\"synapt\") OR (detd:\"desolvated\"ab:\"peaks\"clm:\"spectrometry\"clm:\"spectra\") OR (clm:\"maldi\"detd:\"ftms\"clm:\"electrospray\"detd:\"quadrupoles\"ab:\"spectrometry\"detd:\"spectrom\"ab:\"mass\"ti:\"methods\") OR (detd:\"muddiman\"detd:\"gygi\"detd:\"lysc\") OR (detd:\"lumos\"cpc:\"G16B40/10\"detd:\"lysc\") OR (cpc:\"G16B20/00\"ti:\"complex\"ti:\"sample\"ab:\"ion\") OR (ab:\"spectrometry\"clm:\"spectrometry\"ab:\"peptide\"ab:\"mass\"ab:\"detected\"ab:\"improved\") OR (detd:\"picotip\"detd:\"fibrinopeptide\"clm:\"trna\") OR (detd:\"chait\"detd:\"sequest\"detd:\"proteomes\"detd:\"endoproteinase\") OR (detd:\"spectrom\"clm:\"spectrometry\"clm:\"spectra\"ab:\"peptide\"ab:\"mass\"ab:\"data\"ab:\"invention\"ab:\"one\"ab:\"with\") OR (cpc:\"Y10T436/24\"clm:\"labeled\"ab:\"sample\"ab:\"cell\"ab:\"obtained\") OR (detd:\"silac\"detd:\"itraq\"ti:\"multiplexed\"detd:\"iodoacetyl\") OR (clm:\"biomolecules\"clm:\"fragmenting\"clm:\"biomolecule\"clm:\"abundance\"clm:\"desorption\"clm:\"ionization\"clm:\"spectra\"clm:\"peptides\"clm:\"spectrometer\"clm:\"assisted\"clm:\"proteins\") OR (cpc:\"G01N33/6848\"ti:\"spectrometry\"ti:\"mass\"clm:\"peptides\"ab:\"peptide\"ab:\"mass\"ab:\"sample\"ab:\"determining\"ab:\"present\"ab:\"least\") OR (clm:\"fentomole\")`\n\n- solution without magic(39 targets)\n  - `(clm:\"walwffk\") OR (detd:\"gruhler\" detd:\"qstar\") OR (detd:\"hdmse\" detd:\"roepstorff\") OR (detd:\"gillet\" detd:\"electrosprayed\") OR (detd:\"sofeware\" detd:\"3.4.21.8\") OR (ab:\"srm\" clm:\"srm\" cpc:\"G01N33/6848\" clm:\"quantified\") OR (clm:\"lysophatidylinositol\") OR (detd:\"econometrics\" detd:\"silac\") OR (detd:\"sequest\" ab:\"spectrometry\" clm:\"multidimensional\") OR (detd:\"lumos\" cpc:\"G16B40/10\" detd:\"lysc\") OR (detd:\"phosphoproteomes\" detd:\"picotti\") OR (detd:\"pp4\" detd:\"femtomole\") OR (detd:\"c10h9n6o2\") OR (detd:\"mortz\" cpc:\"H01J49/00\") OR (detd:\"aqyneiqgwdhlsllp\") OR (clm:\"proteome\" detd:\"photodissociation\") OR (detd:\"albar\" detd:\"silac\")`",
      "votes": null
    },
    {
      "id": "2935470",
      "postDate": "07/25/2024 09:58:33",
      "content": "<p>Congratulations on winning the 2nd place in this competition. Thanks for sharing your solution details.<br>\nRegarding discussion on \"Magic\": some bugs or design flaws in the code become great features for machine learning!!</p>",
      "rawMarkdown": "Congratulations on winning the 2nd place in this competition. Thanks for sharing your solution details.\nRegarding discussion on \"Magic\": some bugs or design flaws in the code become great features for machine learning!!",
      "votes": null
    },
    {
      "id": "2935479",
      "postDate": "07/25/2024 10:15:08",
      "content": "<p>Congrats! Really good result!</p>",
      "rawMarkdown": "Congrats! Really good result!",
      "votes": null
    },
    {
      "id": "2935733",
      "postDate": "07/25/2024 14:33:38",
      "content": "<p>Congratulations !</p>\n<p>I understand your disappointment, but don't be too hard on yourself. As for me, I wish I had discovered the magic to claim a cashprize !</p>",
      "rawMarkdown": "Congratulations !\n\nI understand your disappointment, but don't be too hard on yourself. As for me, I wish I had discovered the magic to claim a cashprize !",
      "votes": null
    },
    {
      "id": "2936611",
      "postDate": "07/26/2024 10:24:40",
      "content": "<p>That was helpful ! </p>",
      "rawMarkdown": "That was helpful !",
      "votes": null
    },
    {
      "id": "2940833",
      "postDate": "07/30/2024 13:49:41",
      "content": "<p>Congratulations on your impressive 2nd place finish in the USPTO - Explainable AI for Patent Professionals competition! <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a> <br>\nYour solution, which achieved a public LB score of 0.99849 and a private LB score of 0.99872, demonstrates a high level of precision and effectiveness. The innovative application of \"magic\" techniques in query formulation and the comprehensive insights into the test dataset showcase your exceptional skills and deep understanding of the problem.</p>\n<p>Your meticulous methodology, from constructing indices to generating and selecting subqueries, highlights a sophisticated approach to tackling this complex challenge. The choice to implement your solution in C++ for better performance and memory efficiency, combined with your strategic handling of non-targets, reflects a keen grasp of both technical and practical aspects of the task. Thank you for sharing these valuable insights, and congratulations once again on earning your 5th gold medal and contributing significantly to the field.</p>",
      "rawMarkdown": "Congratulations on your impressive 2nd place finish in the USPTO - Explainable AI for Patent Professionals competition! @shunrcn \nYour solution, which achieved a public LB score of 0.99849 and a private LB score of 0.99872, demonstrates a high level of precision and effectiveness. The innovative application of \"magic\" techniques in query formulation and the comprehensive insights into the test dataset showcase your exceptional skills and deep understanding of the problem.\n\nYour meticulous methodology, from constructing indices to generating and selecting subqueries, highlights a sophisticated approach to tackling this complex challenge. The choice to implement your solution in C++ for better performance and memory efficiency, combined with your strategic handling of non-targets, reflects a keen grasp of both technical and practical aspects of the task. Thank you for sharing these valuable insights, and congratulations once again on earning your 5th gold medal and contributing significantly to the field.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2935470,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/25/2024 09:58:33",
      "content": "<p>Congratulations on winning the 2nd place in this competition. Thanks for sharing your solution details.<br>\nRegarding discussion on \"Magic\": some bugs or design flaws in the code become great features for machine learning!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2935479,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "07/25/2024 10:15:08",
      "content": "<p>Congrats! Really good result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2935733,
      "author_name": "vincentschuler",
      "author_url": "",
      "post_date": "07/25/2024 14:33:38",
      "content": "<p>Congratulations !</p>\n<p>I understand your disappointment, but don't be too hard on yourself. As for me, I wish I had discovered the magic to claim a cashprize !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2936611,
      "author_name": "aadiar",
      "author_url": "",
      "post_date": "07/26/2024 10:24:40",
      "content": "<p>That was helpful ! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2940833,
      "author_name": "",
      "author_url": "",
      "post_date": "07/30/2024 13:49:41",
      "content": "<p>Congratulations on your impressive 2nd place finish in the USPTO - Explainable AI for Patent Professionals competition! <a href=\"https://www.kaggle.com/shunrcn\" target=\"_blank\">@shunrcn</a> <br>\nYour solution, which achieved a public LB score of 0.99849 and a private LB score of 0.99872, demonstrates a high level of precision and effectiveness. The innovative application of \"magic\" techniques in query formulation and the comprehensive insights into the test dataset showcase your exceptional skills and deep understanding of the problem.</p>\n<p>Your meticulous methodology, from constructing indices to generating and selecting subqueries, highlights a sophisticated approach to tackling this complex challenge. The choice to implement your solution in C++ for better performance and memory efficiency, combined with your strategic handling of non-targets, reflects a keen grasp of both technical and practical aspects of the task. Thank you for sharing these valuable insights, and congratulations once again on earning your 5th gold medal and contributing significantly to the field.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935414": "Thanks to the host and congratulations to the winners.\nI'm honestly very disappointed to have missed out on 1st place, but I'm glad to win my first prize and 5th gold medal on this competition!\n\n## Summary\n\n|  | CV | Public LB | Private LB |\n| --- | --- | --- | --- |\n| solution with magic | 0.99934 | 0.99849 | 0.99872 |\n| solution without magic | 0.89992 | 0.90295 | 0.90427 |\n\n## Magic\n\nThere are two (maybe unnoticed by the host) magics in this competition, and I used the former in my solution.\n\n1. no whitespace AND\n  `ti:\"abcd\"ab:\"efgh\"clm:\"ijkl\"`\n2. no whitespace n-gram\n  `ti:\"abcd/efgh/ijkl\"` (there are several possible delimiters other than the slash)\n\nIf you notice one of the two magics, you can get a score above 0.8 even with a simple solution like identifying 1 patent in 1 subquery.\nWithout magic, the best submitted LB was 0.90. However, since I found magic, I focused on improving the solution using it, and I think there is much room for improvement in the solution without magic.\n\n## Insights on Metrics and Test Dataset\n\n-  I suspect that the test_index contains non-targets that are similar to target intentionally. This can be guessed from the fact that when you submit a solution using a query that may include several non-targets, you get a large CV-LB gap, e.g., CV:0.8 vs. LB:0.5.\n-  According to the Whoosh Docs, the order of query results is ordered using TF-IDF scores, and it is difficult to control this order.\n-  This competition's metric has the nature that query results that do not include non-targets will have significantly higher scores. For example, the expected value of metric for 25 targets and 25 non-targets is 0.50 (if the order is random as said in previous discussion), but for 25 targets and 0 non-targets it is 0.84.\n-  From the above discussion, we can see that the solution to construct a query that always has zero non-targets is good.\n-  To achieve this, we can aim for a solution like \"construct an index with data from all patents, and make a query that is determined to contain no non-targets\". However, it is difficult to build Whoosh with all patent data, so it is necessary to build my own fast search algorithm.\n\n## Validation\n\n- I obtained a high correlation with LB by constructing the following CV\n    - Randomly select 2,500 rows from 1975 or later where title/abstract/claims/description are non-null.\n    - Based on the above discussion, limit queries to those that do not include non-targets and absorb the CV/LB gap caused by intentional addition of non-targets in LB.\n\n## Solution\n\n- In the final solution, the query looks like this\n    - `query = (subquery1 OR subquery2 OR …) NOT (negative-subquery1 OR negative-subquery2 OR …)`\n    - Here, subquery and negative-subquery is several words connected by AND.\n    - Negative-subquery is used to cancel out the non-targets that appear in subquery. Using negative-subquery improved the CV by about 0.0002.\n\n- The process is divided into three main parts: \"constructing index\", \"generating subquery candidates\", and \"selecting a subquery to use in a query\".\n    - The \"constructing index\" is done in advance, and the rest of the processing is done for each of the 2500 rows.\n    - For time and memory efficiency, all of the following solution methods were implemented in C++.\n        - Compared to the Python solution, it is about 5 times faster in time and 3 times more efficient in memory.\n\n- Below is an illustration of the entire process.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fe49dfd63ae109f96c6a25abebfcf68b2%2F2024-07-30%20182835.png?generation=1722331846448753&alt=media)\n\n- **Constructing index**\n    - Scan all patent sentences for cpc/title/abstract/claims/description and make a two-dimensional variable-length list from each patent to word, and from each word to each patent.\n    - Words and patents are not treated as strings, but are all converted to integer IDs in advance (for memory efficiency).\n    - Since the data for claims and descriptions is so large that it cannot be stored in RAM, I deleted words with high frequency of occurrence and left words with low frequency of occurrence, thus reducing the data volume to 13% for claims and 1.5% for descriptions.\n    - C++ programs will use about 25GB of RAM, which is enough to satisfy the Notebook's 30GB limit.\n\n- **Generating subquery candidates**\n    - Create subquery candidates corresponding to 50 single target and 50*49/2=1225 target pairs.\n    - Create a base word set as follows.\n         - For a single target, the base word set is all the words it contains.\n         - For a target pair, the base word set is all the words common to both targets.\n    - If a subquery is simply a set of base words connected by AND, the query will take a long time to execute.\n         - Since the intersection set calculation takes min(len(A), len(B)).\n    - Therefore, it is faster to sort in ascending order of the size of the patent corresponding to each word in advance, and to compute the intersection set in that order.\n    - Also, the computation of the  intersection set can be speeded up by terminating the computation as soon as non-targets are removed from the set. This also improves the score, since it is sometimes possible to retrieve targets larger than two.\n    - If the set size is too large at the first computation of the intersection set, give up.\n    - If non-targets remain after a certain number of words, check if it is possible to cancel non-targets using negative-subquery, and if that is not possible, give up.\n    - In this way, the number and order of words to construct a subquery can be optimized, and the results of executing that subquery can be retrieved at the same time.\n\n- **Selecting a subquery to use in a query**\n    - Repeat the following to greedily add subqueries until all 50 targets are covered or the 25 subquery limits are used up.\n        - Score all subquery candidates and select the highest scoring subquery.\n        - The score function is as follows.\n             - $$score(\\text{subquery}) = (\\text{Number of new targets that can be covered by selecting this subquery}) - \\sum\\limits_{\\text{targets covered by this subquery}} (\\text{Number of times the target appears in all subquery candidates}) * 10^{-9}$$\n        - The second term is used for tie-breaking between targets with the same number of new targets that can be covered. The more difficult a target is to cover (fewer candidates), the higher the score when it is covered.\n    - In the final submission, this greedy method was improved very slightly using beam search (CV+0.00003).\n\n- Histogram of the number of targets that could be retrieved in the final sub is as follows.\n    - 91% of the data gave a perfect score (50 targets)\n    - 99% of the data gave more than 45 targets\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3026173%2Fcbe617f06b8231ec355b94ea87d1e652%2F2024-07-25%20172310.png?generation=1721895803039500&alt=media)\n\n## Query Example\n\n- solution with magic(50 targets)\n  - `(ab:\"srm\"clm:\"srm\"cpc:\"G01N33/6848\"clm:\"quantified\") OR (detd:\"collisionally\"cpc:\"G01N33/6848\"detd:\"collisional\"ab:\"spectrometry\"detd:\"spectrom\"clm:\"spectrometry\"clm:\"tandem\"ab:\"proteins\") OR (detd:\"higgs\"ab:\"spectrometry\"clm:\"peptides\"clm:\"ms\"ab:\"proteins\"clm:\"proteins\"ab:\"protein\") OR (detd:\"xics\"detd:\"silac\"detd:\"itraq\"detd:\"xic\"detd:\"sequest\"ab:\"quantifying\") OR (detd:\"massbank\"detd:\"metlin\"detd:\"hmdb\"detd:\"synapt\") OR (detd:\"desolvated\"ab:\"peaks\"clm:\"spectrometry\"clm:\"spectra\") OR (clm:\"maldi\"detd:\"ftms\"clm:\"electrospray\"detd:\"quadrupoles\"ab:\"spectrometry\"detd:\"spectrom\"ab:\"mass\"ti:\"methods\") OR (detd:\"muddiman\"detd:\"gygi\"detd:\"lysc\") OR (detd:\"lumos\"cpc:\"G16B40/10\"detd:\"lysc\") OR (cpc:\"G16B20/00\"ti:\"complex\"ti:\"sample\"ab:\"ion\") OR (ab:\"spectrometry\"clm:\"spectrometry\"ab:\"peptide\"ab:\"mass\"ab:\"detected\"ab:\"improved\") OR (detd:\"picotip\"detd:\"fibrinopeptide\"clm:\"trna\") OR (detd:\"chait\"detd:\"sequest\"detd:\"proteomes\"detd:\"endoproteinase\") OR (detd:\"spectrom\"clm:\"spectrometry\"clm:\"spectra\"ab:\"peptide\"ab:\"mass\"ab:\"data\"ab:\"invention\"ab:\"one\"ab:\"with\") OR (cpc:\"Y10T436/24\"clm:\"labeled\"ab:\"sample\"ab:\"cell\"ab:\"obtained\") OR (detd:\"silac\"detd:\"itraq\"ti:\"multiplexed\"detd:\"iodoacetyl\") OR (clm:\"biomolecules\"clm:\"fragmenting\"clm:\"biomolecule\"clm:\"abundance\"clm:\"desorption\"clm:\"ionization\"clm:\"spectra\"clm:\"peptides\"clm:\"spectrometer\"clm:\"assisted\"clm:\"proteins\") OR (cpc:\"G01N33/6848\"ti:\"spectrometry\"ti:\"mass\"clm:\"peptides\"ab:\"peptide\"ab:\"mass\"ab:\"sample\"ab:\"determining\"ab:\"present\"ab:\"least\") OR (clm:\"fentomole\")`\n\n- solution without magic(39 targets)\n  - `(clm:\"walwffk\") OR (detd:\"gruhler\" detd:\"qstar\") OR (detd:\"hdmse\" detd:\"roepstorff\") OR (detd:\"gillet\" detd:\"electrosprayed\") OR (detd:\"sofeware\" detd:\"3.4.21.8\") OR (ab:\"srm\" clm:\"srm\" cpc:\"G01N33/6848\" clm:\"quantified\") OR (clm:\"lysophatidylinositol\") OR (detd:\"econometrics\" detd:\"silac\") OR (detd:\"sequest\" ab:\"spectrometry\" clm:\"multidimensional\") OR (detd:\"lumos\" cpc:\"G16B40/10\" detd:\"lysc\") OR (detd:\"phosphoproteomes\" detd:\"picotti\") OR (detd:\"pp4\" detd:\"femtomole\") OR (detd:\"c10h9n6o2\") OR (detd:\"mortz\" cpc:\"H01J49/00\") OR (detd:\"aqyneiqgwdhlsllp\") OR (clm:\"proteome\" detd:\"photodissociation\") OR (detd:\"albar\" detd:\"silac\")`",
    "2935470": "Congratulations on winning the 2nd place in this competition. Thanks for sharing your solution details.\nRegarding discussion on \"Magic\": some bugs or design flaws in the code become great features for machine learning!!",
    "2935479": "Congrats! Really good result!",
    "2935733": "Congratulations !\n\nI understand your disappointment, but don't be too hard on yourself. As for me, I wish I had discovered the magic to claim a cashprize !",
    "2936611": "That was helpful !",
    "2940833": "Congratulations on your impressive 2nd place finish in the USPTO - Explainable AI for Patent Professionals competition! @shunrcn \nYour solution, which achieved a public LB score of 0.99849 and a private LB score of 0.99872, demonstrates a high level of precision and effectiveness. The innovative application of \"magic\" techniques in query formulation and the comprehensive insights into the test dataset showcase your exceptional skills and deep understanding of the problem.\n\nYour meticulous methodology, from constructing indices to generating and selecting subqueries, highlights a sophisticated approach to tackling this complex challenge. The choice to implement your solution in C++ for better performance and memory efficiency, combined with your strategic handling of non-targets, reflects a keen grasp of both technical and practical aspects of the task. Thank you for sharing these valuable insights, and congratulations once again on earning your 5th gold medal and contributing significantly to the field."
  },
  "source": "meta"
}