{
  "id": 519826,
  "title": "2 data cleaning checks",
  "url": "/competitions/uspto-explainable-ai/discussion/519826",
  "author_name": "",
  "post_date": "2024-07-13T05:55:03.020011600Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>There are more than 10 million unique genetic sequences across all patent text, they average around 35 characters in length and can exceed 1000 characters. Imo they are not useful as search terms and take up unnecessary space and processing time if you are building a vocabulary of low frequency terms. They can be filtered using:</p>\n<pre><code>dna_pattern = re.() \nrna_pattern = re.() \n</code></pre>\n<p>There are also a lot of mangled/garbage words in the patent text. Many don't have a discernible pattern. Some have non-alphanumeric characters and some do not. A portion can be filtered by setting a max (and min) threshold on the Shannon entropy of words (Shannon entropy has a dependency on word length as well, so the thresholds set should be specific to word length ranges):</p>\n<pre><code> ():\n    freq = Counter(word)\n    total_chars = (word)\n     {char: count / total_chars  char, count  freq.items()}\n\n ():\n    freq = character_frequency(word)\n    entropy = -(p * math.log2(p)  p  freq.values())\n     entropy\n\n\n\n\n\n\n\n</code></pre>",
  "messages": [
    {
      "id": "2919680",
      "postDate": "07/13/2024 05:55:03",
      "content": "<p>There are more than 10 million unique genetic sequences across all patent text, they average around 35 characters in length and can exceed 1000 characters. Imo they are not useful as search terms and take up unnecessary space and processing time if you are building a vocabulary of low frequency terms. They can be filtered using:</p>\n<pre><code>dna_pattern = re.() \nrna_pattern = re.() \n</code></pre>\n<p>There are also a lot of mangled/garbage words in the patent text. Many don't have a discernible pattern. Some have non-alphanumeric characters and some do not. A portion can be filtered by setting a max (and min) threshold on the Shannon entropy of words (Shannon entropy has a dependency on word length as well, so the thresholds set should be specific to word length ranges):</p>\n<pre><code> ():\n    freq = Counter(word)\n    total_chars = (word)\n     {char: count / total_chars  char, count  freq.items()}\n\n ():\n    freq = character_frequency(word)\n    entropy = -(p * math.log2(p)  p  freq.values())\n     entropy\n\n\n\n\n\n\n\n</code></pre>",
      "rawMarkdown": "There are more than 10 million unique genetic sequences across all patent text, they average around 35 characters in length and can exceed 1000 characters. Imo they are not useful as search terms and take up unnecessary space and processing time if you are building a vocabulary of low frequency terms. They can be filtered using:\n\n```python\ndna_pattern = re.compile(r'^[actg\\d]+$') # e.g. 'cttaaatgtc'\nrna_pattern = re.compile(r'^[gacu\\d]+$') # e.g. 'auaaagcuagauaaccgaaagu'\n```\n\nThere are also a lot of mangled/garbage words in the patent text. Many don't have a discernible pattern. Some have non-alphanumeric characters and some do not. A portion can be filtered by setting a max (and min) threshold on the Shannon entropy of words (Shannon entropy has a dependency on word length as well, so the thresholds set should be specific to word length ranges):\n\n```python\ndef character_frequency(word):\n    freq = Counter(word)\n    total_chars = len(word)\n    return {char: count / total_chars for char, count in freq.items()}\n\ndef shannon_entropy(word):\n    freq = character_frequency(word)\n    entropy = -sum(p * math.log2(p) for p in freq.values())\n    return entropy\n\n# some examples\n# ('yyy.yyy.yyy.yyy', 0.7219280948873623)\n# ('geoservices', 2.8453509366224368)\n# ('oxalodihydrazide', 3.327819531114783)\n# ('p6g4v11n35a.chosd9', 4.058813890331201)\n# ('2ixgen8v75ws1h9aiy0', 4.142664355548846)\n```",
      "votes": null
    },
    {
      "id": "2920393",
      "postDate": "07/13/2024 16:09:18",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "2931727",
      "postDate": "07/22/2024 09:38:32",
      "content": "<p>Thanks for useful information. I incorporated these filters and could not see any improvement in score. </p>",
      "rawMarkdown": "Thanks for useful information. I incorporated these filters and could not see any improvement in score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2920393,
      "author_name": "keakohv",
      "author_url": "",
      "post_date": "07/13/2024 16:09:18",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2931727,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/22/2024 09:38:32",
      "content": "<p>Thanks for useful information. I incorporated these filters and could not see any improvement in score. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2919680": "There are more than 10 million unique genetic sequences across all patent text, they average around 35 characters in length and can exceed 1000 characters. Imo they are not useful as search terms and take up unnecessary space and processing time if you are building a vocabulary of low frequency terms. They can be filtered using:\n\n```python\ndna_pattern = re.compile(r'^[actg\\d]+$') # e.g. 'cttaaatgtc'\nrna_pattern = re.compile(r'^[gacu\\d]+$') # e.g. 'auaaagcuagauaaccgaaagu'\n```\n\nThere are also a lot of mangled/garbage words in the patent text. Many don't have a discernible pattern. Some have non-alphanumeric characters and some do not. A portion can be filtered by setting a max (and min) threshold on the Shannon entropy of words (Shannon entropy has a dependency on word length as well, so the thresholds set should be specific to word length ranges):\n\n```python\ndef character_frequency(word):\n    freq = Counter(word)\n    total_chars = len(word)\n    return {char: count / total_chars for char, count in freq.items()}\n\ndef shannon_entropy(word):\n    freq = character_frequency(word)\n    entropy = -sum(p * math.log2(p) for p in freq.values())\n    return entropy\n\n# some examples\n# ('yyy.yyy.yyy.yyy', 0.7219280948873623)\n# ('geoservices', 2.8453509366224368)\n# ('oxalodihydrazide', 3.327819531114783)\n# ('p6g4v11n35a.chosd9', 4.058813890331201)\n# ('2ixgen8v75ws1h9aiy0', 4.142664355548846)\n```",
    "2920393": "Thank you for sharing!",
    "2931727": "Thanks for useful information. I incorporated these filters and could not see any improvement in score."
  },
  "source": "meta"
}