{
  "id": 522199,
  "title": "The \"Magic\" !",
  "url": "/competitions/uspto-explainable-ai/discussion/522199",
  "author_name": "Theo Viel",
  "post_date": "2024-07-25T00:01:38.051000",
  "votes": 46,
  "comment_count": 14,
  "views": 0,
  "content": "<p><strong>TLDR:</strong> You can create queries as long as you want by using a \"fake\" space character. Whoosh's preprocessing will replace this character by a whitespace, whereas the <code>count_query_tokens</code> function will not !</p>\n<p>Sometimes a code example is better than a long explanation:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/theoviel/uspto-magic-with-cudf-pandas-lb-0-8-in-1min\" target=\"_blank\">https://www.kaggle.com/code/theoviel/uspto-magic-with-cudf-pandas-lb-0-8-in-1min</a></p>\n</blockquote>\n<h4>Overall idea</h4>\n<p>Let's say we have two publications ( <code>publication_1</code> and <code>publication_2</code>) with titles <code>title_1 =  t1_1 t1_2 t1_3 ...</code>and <code>title_2 =  t2_1 t2_2 t2_3 ...</code>.<br>\nThe easiest way to query both publications would be to do : <br>\n<code>query = ti:\"title_1\" OR ti:\"title_2\"</code><br>\nThe issue is that titles can be several words long, and with the token limit of 50, you cannot really query enough titles to get a reasonable score.<br>\nIf only title_1 could count as only one token, that would mean we could query 25 titles, and get a score of 0.8+ effortlessly ! </p>\n<h4>The magic</h4>\n<p>The count token function uses spaces:</p>\n<pre><code> ():\n     ([i  i  re.split(, query)  i])\n</code></pre>\n<p>However, whoosh uses a more sophisticated preprocessing. If we can find a character than is replaced by a whitespace by whoosh, but not by the <code>count_query_tokens</code> regex, then we can query a title for the cost of one token  !</p>\n<p>And it turns out, such special character exists: <code>~</code> is an example. This query :</p>\n<pre><code> = ti: OR ti:\n</code></pre>\n<p>Costs 3 tokens and will match the same document as </p>\n<pre><code> = ti: OR ti:\n</code></pre>\n<p>Using this trick to query 25 exact titles alone achieves LB 0.8 which is almost in the gold zone.</p>\n<p>Using <a href=\"https://rapids.ai/cudf-pandas/\" target=\"_blank\">cudf-pandas</a> to process everything on GPU makes the code super-fast. The notebook shared above runs in 1 minute ! Unfortunately there is no efficiency track in this competition 😄</p>\n<p>We used slightly fancier heuristics to reach 7th place, but I'm sure top teams found even better hacks !</p>",
  "messages": [
    {
      "id": 2935053,
      "postDate": "2024-07-25T00:01:38.050Z",
      "content": "<p><strong>TLDR:</strong> You can create queries as long as you want by using a \"fake\" space character. Whoosh's preprocessing will replace this character by a whitespace, whereas the <code>count_query_tokens</code> function will not !</p>\n<p>Sometimes a code example is better than a long explanation:</p>\n<blockquote>\n  <p><a href=\"https://www.kaggle.com/code/theoviel/uspto-magic-with-cudf-pandas-lb-0-8-in-1min\" target=\"_blank\">https://www.kaggle.com/code/theoviel/uspto-magic-with-cudf-pandas-lb-0-8-in-1min</a></p>\n</blockquote>\n<h4>Overall idea</h4>\n<p>Let's say we have two publications ( <code>publication_1</code> and <code>publication_2</code>) with titles <code>title_1 =  t1_1 t1_2 t1_3 ...</code>and <code>title_2 =  t2_1 t2_2 t2_3 ...</code>.<br>\nThe easiest way to query both publications would be to do : <br>\n<code>query = ti:\"title_1\" OR ti:\"title_2\"</code><br>\nThe issue is that titles can be several words long, and with the token limit of 50, you cannot really query enough titles to get a reasonable score.<br>\nIf only title_1 could count as only one token, that would mean we could query 25 titles, and get a score of 0.8+ effortlessly ! </p>\n<h4>The magic</h4>\n<p>The count token function uses spaces:</p>\n<pre><code> ():\n     ([i  i  re.split(, query)  i])\n</code></pre>\n<p>However, whoosh uses a more sophisticated preprocessing. If we can find a character than is replaced by a whitespace by whoosh, but not by the <code>count_query_tokens</code> regex, then we can query a title for the cost of one token  !</p>\n<p>And it turns out, such special character exists: <code>~</code> is an example. This query :</p>\n<pre><code> = ti: OR ti:\n</code></pre>\n<p>Costs 3 tokens and will match the same document as </p>\n<pre><code> = ti: OR ti:\n</code></pre>\n<p>Using this trick to query 25 exact titles alone achieves LB 0.8 which is almost in the gold zone.</p>\n<p>Using <a href=\"https://rapids.ai/cudf-pandas/\" target=\"_blank\">cudf-pandas</a> to process everything on GPU makes the code super-fast. The notebook shared above runs in 1 minute ! Unfortunately there is no efficiency track in this competition 😄</p>\n<p>We used slightly fancier heuristics to reach 7th place, but I'm sure top teams found even better hacks !</p>",
      "rawMarkdown": "**TLDR:** You can create queries as long as you want by using a \"fake\" space character. Whoosh's preprocessing will replace this character by a whitespace, whereas the `count_query_tokens` function will not !\n\nSometimes a code example is better than a long explanation:\n>https://www.kaggle.com/code/theoviel/uspto-magic-with-cudf-pandas-lb-0-8-in-1min\n\n#### Overall idea\n\nLet's say we have two publications ( `publication_1` and `publication_2`) with titles `title_1 =  t1_1 t1_2 t1_3 ...`and `title_2 =  t2_1 t2_2 t2_3 ...`.\nThe easiest way to query both publications would be to do : \n`query = ti:\"title_1\" OR ti:\"title_2\"`\nThe issue is that titles can be several words long, and with the token limit of 50, you cannot really query enough titles to get a reasonable score.\nIf only title_1 could count as only one token, that would mean we could query 25 titles, and get a score of 0.8+ effortlessly ! \n\n#### The magic\n\nThe count token function uses spaces:\n\n```\ndef count_query_tokens(query: str):\n    return len([i for i in re.split('[\\s+()]', query) if i])\n```\n\nHowever, whoosh uses a more sophisticated preprocessing. If we can find a character than is replaced by a whitespace by whoosh, but not by the `count_query_tokens` regex, then we can query a title for the cost of one token  !\n\nAnd it turns out, such special character exists: `~` is an example. This query :\n```\nquery = ti:\"t1_1~t1_2~t1_3~...\" OR ti:\"t2_1~t2_2~t2_3~...\"\n```\nCosts 3 tokens and will match the same document as \n```\nquery = ti:\"t1_1 t1_2 t1_3...\" OR ti:\"t2_1 t2_2 t2_3...\"\n```\n\nUsing this trick to query 25 exact titles alone achieves LB 0.8 which is almost in the gold zone.\n\nUsing [cudf-pandas](https://rapids.ai/cudf-pandas/) to process everything on GPU makes the code super-fast. The notebook shared above runs in 1 minute ! Unfortunately there is no efficiency track in this competition 😄\n\nWe used slightly fancier heuristics to reach 7th place, but I'm sure top teams found even better hacks !",
      "votes": 46
    },
    {
      "id": 2935152,
      "postDate": "2024-07-25T03:08:36.200Z",
      "content": "<p>Congrats! Hilarious that querying exact titles gets one almost to the gold zone… and in one minute. :D What a slap to complex computation-heavy solutions. Nice demonstration of cuDF! :)</p>",
      "rawMarkdown": "Congrats! Hilarious that querying exact titles gets one almost to the gold zone... and in one minute. :D What a slap to complex computation-heavy solutions. Nice demonstration of cuDF! :)",
      "votes": 3
    },
    {
      "id": 2935055,
      "postDate": "2024-07-25T00:07:12.813Z",
      "content": "<p>25 guessed titles gets a 0.8+ score? Thats cool, the tilda solution is neat 😁, wonder what insights organizers were hoping to get from kaggle solutions, suppose it was expected NNs or something like that will be used, but algo wins today!</p>",
      "rawMarkdown": "25 guessed titles gets a 0.8+ score? Thats cool, the tilda solution is neat 😁, wonder what insights organizers were hoping to get from kaggle solutions, suppose it was expected NNs or something like that will be used, but algo wins today!",
      "votes": 3
    },
    {
      "id": 2936005,
      "postDate": "2024-07-25T17:30:10.307Z",
      "content": "<p>Not sure about this, but why this approach is considered praiseworthy by the community? From the competition overview it was clear that it was aimed at interpretebility (\"Capable solutions to this challenge will enable patent professionals to use AI-powered search capabilities with increased confidence by providing them the ability to interpret the results in a familiar language and syntax.\"), and not at simple retrieval of all the elements by \"magic\". Strange to me that so many participants discovered it and yet no one started the discussion to let organizers know about such a \"magic\" (basically, a bag in the <code>count_query_tokens()</code>) . Again sorry for the comment, I am new to Kaggle and just wanted to understand the community reaction, which was in contrast with my thoughts.</p>",
      "rawMarkdown": "Not sure about this, but why this approach is considered praiseworthy by the community? From the competition overview it was clear that it was aimed at interpretebility (\"Capable solutions to this challenge will enable patent professionals to use AI-powered search capabilities with increased confidence by providing them the ability to interpret the results in a familiar language and syntax.\"), and not at simple retrieval of all the elements by \"magic\". Strange to me that so many participants discovered it and yet no one started the discussion to let organizers know about such a \"magic\" (basically, a bag in the `count_query_tokens()`) . Again sorry for the comment, I am new to Kaggle and just wanted to understand the community reaction, which was in contrast with my thoughts.",
      "votes": 3,
      "replies": [
        {
          "id": 2936021,
          "postDate": "2024-07-25T17:42:35.977Z",
          "content": "<p>I agree with you, I was too disappointed to discover this secret. But I can still see many interesting ideas used by participants, many of them getting really high scores even without this trick. So as far as we consider only insights and usefullness of the competition for the purpose of the explainable AI, I don't think this bug spoils the competition entirely.</p>",
          "rawMarkdown": "I agree with you, I was too disappointed to discover this secret. But I can still see many interesting ideas used by participants, many of them getting really high scores even without this trick. So as far as we consider only insights and usefullness of the competition for the purpose of the explainable AI, I don't think this bug spoils the competition entirely.",
          "votes": 2
        },
        {
          "id": 2936217,
          "postDate": "2024-07-25T23:16:49.693Z",
          "content": "<p>Well it’s kaggle, you squeeze out metrics here. It doesn’t mean that it has no value, there are many things to learn here in the process, but it’s about getting best metrics (if it’s about winning competitions ofc)</p>",
          "rawMarkdown": "Well it’s kaggle, you squeeze out metrics here. It doesn’t mean that it has no value, there are many things to learn here in the process, but it’s about getting best metrics (if it’s about winning competitions ofc)",
          "replies": [
            {
              "id": 2936511,
              "postDate": "2024-07-26T08:11:38.857Z",
              "content": "<p>I agree that it feels like the excessive use of \"magic\" undermines the goal of the competition but fortunately several top competitors also had high-scoring non-magic or a-little-magic solutions and seems like they put a lot of work into their solutions…. Some tips-and-tricks were shared on the forum during the competition (like AND being the default operator) but I guess the real magic was found too late in the competition so no one had the motivation to share.</p>",
              "rawMarkdown": "I agree that it feels like the excessive use of \"magic\" undermines the goal of the competition but fortunately several top competitors also had high-scoring non-magic or a-little-magic solutions and seems like they put a lot of work into their solutions.... Some tips-and-tricks were shared on the forum during the competition (like AND being the default operator) but I guess the real magic was found too late in the competition so no one had the motivation to share.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2936607,
      "postDate": "2024-07-26T10:22:17.673Z",
      "content": "<p>That was an interesting read ! and the title was 🌟</p>",
      "rawMarkdown": "That was an interesting read ! and the title was 🌟",
      "votes": 1
    },
    {
      "id": 2935190,
      "postDate": "2024-07-25T03:53:06.607Z",
      "content": "<p>I used the question mark '?' for my magic symbol. It makes sense that it would work to match the spaces and the whole string would be counted as a single token.</p>",
      "rawMarkdown": "I used the question mark '?' for my magic symbol. It makes sense that it would work to match the spaces and the whole string would be counted as a single token.",
      "votes": 1
    },
    {
      "id": 2935054,
      "postDate": "2024-07-25T00:04:24.060Z",
      "content": "<p>Amazing <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> congrats for the 7th position!</p>",
      "rawMarkdown": "Amazing @theoviel congrats for the 7th position!",
      "votes": 1
    },
    {
      "id": 2937442,
      "postDate": "2024-07-27T05:12:05.807Z",
      "content": "<p>Very interesting approach!</p>",
      "rawMarkdown": "Very interesting approach!"
    },
    {
      "id": 2935433,
      "postDate": "2024-07-25T08:41:24.333Z",
      "content": "<p>Impressive. It is the great thought to implement such logic.</p>",
      "rawMarkdown": "Impressive. It is the great thought to implement such logic."
    },
    {
      "id": 2935157,
      "postDate": "2024-07-25T03:13:49.577Z",
      "content": "<p>This is really wonderful, I have been to optimize the choice of keywords, thank you for sharing☺️☺️☺️</p>",
      "rawMarkdown": "This is really wonderful, I have been to optimize the choice of keywords, thank you for sharing☺️☺️☺️"
    },
    {
      "id": 2935123,
      "postDate": "2024-07-25T01:57:41.550Z",
      "content": "<p>Impressive!</p>",
      "rawMarkdown": "Impressive!"
    },
    {
      "id": 2938295,
      "postDate": "2024-07-28T01:35:34.847Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2935152,
      "author_name": "Kea Kohv",
      "author_url": "",
      "post_date": "2024-07-25T03:08:36.200000",
      "content": "<p>Congrats! Hilarious that querying exact titles gets one almost to the gold zone… and in one minute. :D What a slap to complex computation-heavy solutions. Nice demonstration of cuDF! :)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2935055,
      "author_name": "opamusora (Ivan Viakhirev)",
      "author_url": "",
      "post_date": "2024-07-25T00:07:12.813000",
      "content": "<p>25 guessed titles gets a 0.8+ score? Thats cool, the tilda solution is neat 😁, wonder what insights organizers were hoping to get from kaggle solutions, suppose it was expected NNs or something like that will be used, but algo wins today!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2936005,
      "author_name": "Anastas Akopian",
      "author_url": "",
      "post_date": "2024-07-25T17:30:10.307000",
      "content": "<p>Not sure about this, but why this approach is considered praiseworthy by the community? From the competition overview it was clear that it was aimed at interpretebility (\"Capable solutions to this challenge will enable patent professionals to use AI-powered search capabilities with increased confidence by providing them the ability to interpret the results in a familiar language and syntax.\"), and not at simple retrieval of all the elements by \"magic\". Strange to me that so many participants discovered it and yet no one started the discussion to let organizers know about such a \"magic\" (basically, a bag in the <code>count_query_tokens()</code>) . Again sorry for the comment, I am new to Kaggle and just wanted to understand the community reaction, which was in contrast with my thoughts.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2936021,
          "author_name": "Pavel Kazlou",
          "author_url": "",
          "post_date": "2024-07-25T17:42:35.977000",
          "content": "<p>I agree with you, I was too disappointed to discover this secret. But I can still see many interesting ideas used by participants, many of them getting really high scores even without this trick. So as far as we consider only insights and usefullness of the competition for the purpose of the explainable AI, I don't think this bug spoils the competition entirely.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2936217,
          "author_name": "opamusora (Ivan Viakhirev)",
          "author_url": "",
          "post_date": "2024-07-25T23:16:49.693000",
          "content": "<p>Well it’s kaggle, you squeeze out metrics here. It doesn’t mean that it has no value, there are many things to learn here in the process, but it’s about getting best metrics (if it’s about winning competitions ofc)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2936511,
              "author_name": "Kea Kohv",
              "author_url": "",
              "post_date": "2024-07-26T08:11:38.857000",
              "content": "<p>I agree that it feels like the excessive use of \"magic\" undermines the goal of the competition but fortunately several top competitors also had high-scoring non-magic or a-little-magic solutions and seems like they put a lot of work into their solutions…. Some tips-and-tricks were shared on the forum during the competition (like AND being the default operator) but I guess the real magic was found too late in the competition so no one had the motivation to share.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2936607,
      "author_name": "aadiAR",
      "author_url": "",
      "post_date": "2024-07-26T10:22:17.673000",
      "content": "<p>That was an interesting read ! and the title was 🌟</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2935190,
      "author_name": "Oleg Kokorin",
      "author_url": "",
      "post_date": "2024-07-25T03:53:06.607000",
      "content": "<p>I used the question mark '?' for my magic symbol. It makes sense that it would work to match the spaces and the whole string would be counted as a single token.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2935054,
      "author_name": "Octavi Grau",
      "author_url": "",
      "post_date": "2024-07-25T00:04:24.060000",
      "content": "<p>Amazing <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> congrats for the 7th position!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2937442,
      "author_name": "DenisIN4",
      "author_url": "",
      "post_date": "2024-07-27T05:12:05.807000",
      "content": "<p>Very interesting approach!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2935433,
      "author_name": "Devang Garg",
      "author_url": "",
      "post_date": "2024-07-25T08:41:24.333000",
      "content": "<p>Impressive. It is the great thought to implement such logic.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2935157,
      "author_name": "Alannikos",
      "author_url": "",
      "post_date": "2024-07-25T03:13:49.577000",
      "content": "<p>This is really wonderful, I have been to optimize the choice of keywords, thank you for sharing☺️☺️☺️</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2935123,
      "author_name": "LIGHT",
      "author_url": "",
      "post_date": "2024-07-25T01:57:41.550000",
      "content": "<p>Impressive!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2938295,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-28T01:35:34.847000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935053": "**TLDR:** You can create queries as long as you want by using a \"fake\" space character. Whoosh's preprocessing will replace this character by a whitespace, whereas the `count_query_tokens` function will not !\n\nSometimes a code example is better than a long explanation:\n>https://www.kaggle.com/code/theoviel/uspto-magic-with-cudf-pandas-lb-0-8-in-1min\n\n#### Overall idea\n\nLet's say we have two publications ( `publication_1` and `publication_2`) with titles `title_1 =  t1_1 t1_2 t1_3 ...`and `title_2 =  t2_1 t2_2 t2_3 ...`.\nThe easiest way to query both publications would be to do : \n`query = ti:\"title_1\" OR ti:\"title_2\"`\nThe issue is that titles can be several words long, and with the token limit of 50, you cannot really query enough titles to get a reasonable score.\nIf only title_1 could count as only one token, that would mean we could query 25 titles, and get a score of 0.8+ effortlessly ! \n\n#### The magic\n\nThe count token function uses spaces:\n\n```\ndef count_query_tokens(query: str):\n    return len([i for i in re.split('[\\s+()]', query) if i])\n```\n\nHowever, whoosh uses a more sophisticated preprocessing. If we can find a character than is replaced by a whitespace by whoosh, but not by the `count_query_tokens` regex, then we can query a title for the cost of one token  !\n\nAnd it turns out, such special character exists: `~` is an example. This query :\n```\nquery = ti:\"t1_1~t1_2~t1_3~...\" OR ti:\"t2_1~t2_2~t2_3~...\"\n```\nCosts 3 tokens and will match the same document as \n```\nquery = ti:\"t1_1 t1_2 t1_3...\" OR ti:\"t2_1 t2_2 t2_3...\"\n```\n\nUsing this trick to query 25 exact titles alone achieves LB 0.8 which is almost in the gold zone.\n\nUsing [cudf-pandas](https://rapids.ai/cudf-pandas/) to process everything on GPU makes the code super-fast. The notebook shared above runs in 1 minute ! Unfortunately there is no efficiency track in this competition 😄\n\nWe used slightly fancier heuristics to reach 7th place, but I'm sure top teams found even better hacks !",
    "2935152": "Congrats! Hilarious that querying exact titles gets one almost to the gold zone... and in one minute. :D What a slap to complex computation-heavy solutions. Nice demonstration of cuDF! :)",
    "2935055": "25 guessed titles gets a 0.8+ score? Thats cool, the tilda solution is neat 😁, wonder what insights organizers were hoping to get from kaggle solutions, suppose it was expected NNs or something like that will be used, but algo wins today!",
    "2936005": "Not sure about this, but why this approach is considered praiseworthy by the community? From the competition overview it was clear that it was aimed at interpretebility (\"Capable solutions to this challenge will enable patent professionals to use AI-powered search capabilities with increased confidence by providing them the ability to interpret the results in a familiar language and syntax.\"), and not at simple retrieval of all the elements by \"magic\". Strange to me that so many participants discovered it and yet no one started the discussion to let organizers know about such a \"magic\" (basically, a bag in the `count_query_tokens()`) . Again sorry for the comment, I am new to Kaggle and just wanted to understand the community reaction, which was in contrast with my thoughts.",
    "2936607": "That was an interesting read ! and the title was 🌟",
    "2935190": "I used the question mark '?' for my magic symbol. It makes sense that it would work to match the spaces and the whole string would be counted as a single token.",
    "2935054": "Amazing @theoviel congrats for the 7th position!",
    "2937442": "Very interesting approach!",
    "2935433": "Impressive. It is the great thought to implement such logic.",
    "2935157": "This is really wonderful, I have been to optimize the choice of keywords, thank you for sharing☺️☺️☺️",
    "2935123": "Impressive!",
    "2938295": ""
  }
}