{
  "id": 519755,
  "title": "Is there any way to see the whoosh index content for a certain document?",
  "url": "/competitions/uspto-explainable-ai/discussion/519755",
  "author_name": "",
  "post_date": "2024-07-12T16:43:55.057976600Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have run into a problem that certain strings are not found by whoosh even though they supposedly should.<br>\nFor example, there is the following text in the patent description:</p>\n<p>Probability<br>\n0.395<br>\n0.0932 <br>\nVariety</p>\n<p>custom_analyzer will produce these tokens: ['probability', 'variety']<br>\nHowever, if I run the query detd:\"probability variety\", the document is not found<br>\nIf I run detd:probability AND detd:variety, the document is found, i.e. it is definitely in the index.</p>\n<p>It looks like the numbers can break the token sequence even though they are not indexed, which is annoying. I want to see the index content to understand what query would work in this case. What is the best way to do it?</p>",
  "messages": [
    {
      "id": "2919047",
      "postDate": "07/12/2024 16:43:55",
      "content": "<p>I have run into a problem that certain strings are not found by whoosh even though they supposedly should.<br>\nFor example, there is the following text in the patent description:</p>\n<p>Probability<br>\n0.395<br>\n0.0932 <br>\nVariety</p>\n<p>custom_analyzer will produce these tokens: ['probability', 'variety']<br>\nHowever, if I run the query detd:\"probability variety\", the document is not found<br>\nIf I run detd:probability AND detd:variety, the document is found, i.e. it is definitely in the index.</p>\n<p>It looks like the numbers can break the token sequence even though they are not indexed, which is annoying. I want to see the index content to understand what query would work in this case. What is the best way to do it?</p>",
      "rawMarkdown": "I have run into a problem that certain strings are not found by whoosh even though they supposedly should.\nFor example, there is the following text in the patent description:\n\nProbability\n0.395\n0.0932 \nVariety\n\ncustom_analyzer will produce these tokens: ['probability', 'variety']\nHowever, if I run the query detd:\"probability variety\", the document is not found\nIf I run detd:probability AND detd:variety, the document is found, i.e. it is definitely in the index.\n\nIt looks like the numbers can break the token sequence even though they are not indexed, which is annoying. I want to see the index content to understand what query would work in this case. What is the best way to do it?",
      "votes": null
    },
    {
      "id": "2919290",
      "postDate": "07/12/2024 19:14:20",
      "content": "<p>How about trying this completely? detd:\"Probability 0.395 0.0932 Variety\"</p>",
      "rawMarkdown": "How about trying this completely? detd:\"Probability 0.395 0.0932 Variety\"",
      "votes": null
    },
    {
      "id": "2919472",
      "postDate": "07/12/2024 22:51:30",
      "content": "<p>It seems possible to obtain the positions by passing the argument <code>positions=True</code> to the analyzer.</p>\n<pre><code> whoosh\n whoosh.analysis\n re\n\nBRS_STOPWORDS = [, , , , , , , , , , , ,\n    , , , , , , , , , , ]\nNUMBER_REGEX = re.()\n (whoosh.analysis.Filter):\n     ():\n         t  tokens:\n              NUMBER_REGEX.(t.text):\n                 t\n\nanalyzer = whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()\n token  analyzer(, positions=):\n    (token.text, token.pos)\n</code></pre>\n<pre><code> \n \n</code></pre>\n<p>For more details: <a href=\"https://whoosh.readthedocs.io/en/latest/analysis.html#token-information-attributes\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/analysis.html#token-information-attributes</a></p>\n<p>It seems that unlike whoosh.analysis.StopFilter, whoosh_utils.NumberFilter does not renumber the position of each token, so the position before excluding numbers is returned.</p>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/analysis.html#renumbering-term-positions\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/analysis.html#renumbering-term-positions</a><br>\n<a href=\"https://github.com/mchaput/whoosh/blob/d9a3fa2a4905e7326c9623c89e6395713c189161/src/whoosh/analysis/filters.py#L307-L312\" target=\"_blank\">https://github.com/mchaput/whoosh/blob/d9a3fa2a4905e7326c9623c89e6395713c189161/src/whoosh/analysis/filters.py#L307-L312</a></p>",
      "rawMarkdown": "It seems possible to obtain the positions by passing the argument `positions=True` to the analyzer.\n\n```py\nimport whoosh\nimport whoosh.analysis\nimport re\n\nBRS_STOPWORDS = ['an', 'are', 'by', 'for', 'if', 'into', 'is', 'no', 'not', 'of', 'on', 'such',\n    'that', 'the', 'their', 'then', 'there', 'these', 'they', 'this', 'to', 'was', 'will']\nNUMBER_REGEX = re.compile(r'^(\\d+|\\d{1,3}(,\\d{3})*)(\\.\\d+)?$')\nclass NumberFilter(whoosh.analysis.Filter):\n    def __call__(self, tokens):\n        for t in tokens:\n            if not NUMBER_REGEX.match(t.text):\n                yield t\n\nanalyzer = whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()\nfor token in analyzer(\"Probability 0.395 0.0932 Variety\", positions=True):\n    print(token.text, token.pos)\n```\n\n```\nprobability 0\nvariety 3\n```\n\nFor more details: https://whoosh.readthedocs.io/en/latest/analysis.html#token-information-attributes\n\nIt seems that unlike whoosh.analysis.StopFilter, whoosh_utils.NumberFilter does not renumber the position of each token, so the position before excluding numbers is returned.\n\nhttps://whoosh.readthedocs.io/en/latest/analysis.html#renumbering-term-positions\nhttps://github.com/mchaput/whoosh/blob/d9a3fa2a4905e7326c9623c89e6395713c189161/src/whoosh/analysis/filters.py#L307-L312",
      "votes": null
    },
    {
      "id": "2920236",
      "postDate": "07/13/2024 14:24:00",
      "content": "<p>Thank you! I found a way to overcome the problem at hand, but your solution is certainly more elegant and precise.</p>",
      "rawMarkdown": "Thank you! I found a way to overcome the problem at hand, but your solution is certainly more elegant and precise.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2919290,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "07/12/2024 19:14:20",
      "content": "<p>How about trying this completely? detd:\"Probability 0.395 0.0932 Variety\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2919472,
      "author_name": "sash2104",
      "author_url": "",
      "post_date": "07/12/2024 22:51:30",
      "content": "<p>It seems possible to obtain the positions by passing the argument <code>positions=True</code> to the analyzer.</p>\n<pre><code> whoosh\n whoosh.analysis\n re\n\nBRS_STOPWORDS = [, , , , , , , , , , , ,\n    , , , , , , , , , , ]\nNUMBER_REGEX = re.()\n (whoosh.analysis.Filter):\n     ():\n         t  tokens:\n              NUMBER_REGEX.(t.text):\n                 t\n\nanalyzer = whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()\n token  analyzer(, positions=):\n    (token.text, token.pos)\n</code></pre>\n<pre><code> \n \n</code></pre>\n<p>For more details: <a href=\"https://whoosh.readthedocs.io/en/latest/analysis.html#token-information-attributes\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/analysis.html#token-information-attributes</a></p>\n<p>It seems that unlike whoosh.analysis.StopFilter, whoosh_utils.NumberFilter does not renumber the position of each token, so the position before excluding numbers is returned.</p>\n<p><a href=\"https://whoosh.readthedocs.io/en/latest/analysis.html#renumbering-term-positions\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/analysis.html#renumbering-term-positions</a><br>\n<a href=\"https://github.com/mchaput/whoosh/blob/d9a3fa2a4905e7326c9623c89e6395713c189161/src/whoosh/analysis/filters.py#L307-L312\" target=\"_blank\">https://github.com/mchaput/whoosh/blob/d9a3fa2a4905e7326c9623c89e6395713c189161/src/whoosh/analysis/filters.py#L307-L312</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2920236,
          "author_name": "apparition",
          "author_url": "",
          "post_date": "07/13/2024 14:24:00",
          "content": "<p>Thank you! I found a way to overcome the problem at hand, but your solution is certainly more elegant and precise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2919047": "I have run into a problem that certain strings are not found by whoosh even though they supposedly should.\nFor example, there is the following text in the patent description:\n\nProbability\n0.395\n0.0932 \nVariety\n\ncustom_analyzer will produce these tokens: ['probability', 'variety']\nHowever, if I run the query detd:\"probability variety\", the document is not found\nIf I run detd:probability AND detd:variety, the document is found, i.e. it is definitely in the index.\n\nIt looks like the numbers can break the token sequence even though they are not indexed, which is annoying. I want to see the index content to understand what query would work in this case. What is the best way to do it?",
    "2919290": "How about trying this completely? detd:\"Probability 0.395 0.0932 Variety\"",
    "2919472": "It seems possible to obtain the positions by passing the argument `positions=True` to the analyzer.\n\n```py\nimport whoosh\nimport whoosh.analysis\nimport re\n\nBRS_STOPWORDS = ['an', 'are', 'by', 'for', 'if', 'into', 'is', 'no', 'not', 'of', 'on', 'such',\n    'that', 'the', 'their', 'then', 'there', 'these', 'they', 'this', 'to', 'was', 'will']\nNUMBER_REGEX = re.compile(r'^(\\d+|\\d{1,3}(,\\d{3})*)(\\.\\d+)?$')\nclass NumberFilter(whoosh.analysis.Filter):\n    def __call__(self, tokens):\n        for t in tokens:\n            if not NUMBER_REGEX.match(t.text):\n                yield t\n\nanalyzer = whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()\nfor token in analyzer(\"Probability 0.395 0.0932 Variety\", positions=True):\n    print(token.text, token.pos)\n```\n\n```\nprobability 0\nvariety 3\n```\n\nFor more details: https://whoosh.readthedocs.io/en/latest/analysis.html#token-information-attributes\n\nIt seems that unlike whoosh.analysis.StopFilter, whoosh_utils.NumberFilter does not renumber the position of each token, so the position before excluding numbers is returned.\n\nhttps://whoosh.readthedocs.io/en/latest/analysis.html#renumbering-term-positions\nhttps://github.com/mchaput/whoosh/blob/d9a3fa2a4905e7326c9623c89e6395713c189161/src/whoosh/analysis/filters.py#L307-L312",
    "2920236": "Thank you! I found a way to overcome the problem at hand, but your solution is certainly more elegant and precise."
  },
  "source": "meta"
}