{
  "id": 522347,
  "title": "42th place: a very simple solution",
  "url": "/competitions/uspto-explainable-ai/discussion/522347",
  "author_name": "Pavel Kazlou",
  "post_date": "2024-07-25T18:09:32",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The idea of my solution was very simple: using document frequencies to select the most important tokens and then join top-25 most important tokens with \"OR\" operator.</p>\n<p>Calculating document frequencies for CPC codes was easy and computationally not demanding, so I've made it inline in the submission:</p>\n<pre><code>code_freq = pl.read.explode.value.sort('count', descending=True)  \n\ncode_freq = code_freq.rows\n</code></pre>\n<p>Titles were a little bit more heavy, but still not worth a separate notebook:</p>\n<pre><code>vectorizer = (max_df=, min_df=, binary=True)\n\ntitles_df = pl()\\\n        ()\\\n        ()\n\nword_title_freq = vectorizer(titles_df)\n\nword_title_freq = (\n    (\n        vectorizer(), \n        np(np(word_title_freq(axis=)))\n    )\n)\n</code></pre>\n<p>Having these global document frequencies for CPC codes and words in titles, I just calculated local document frequencies for patents who were neighbors and calculated importance of a token as <code>local_document_frequency / global_document_frequency</code>:</p>\n<pre><code> = [code for codes in row[] for code in codes ]\n = Counter(flat_list_codes)\n = [(f, count / code_freq.get(elem, )) for (elem, count) in codes_counter.items()]\n\n = [title for title in row[] if title]\n = [token.text for title in titles for token in analyzer(title)]\n = Counter(flat_list_title_words)\n = [(f, count / word_title_freq.get(elem, )) for (elem, count) in title_words_counter.items()]\n</code></pre>\n<p>Then just sorted all the tokens by their importance and combined via OR:</p>\n<pre><code>selected_operands = sorted(\n            weighted_codes + weighted_title_words,\n            =lambda x: x[1], \n            =)[:25]\n\nreturn .join(selected_operands)\n</code></pre>\n<p>I've tried experimenting with same approach for pairs of codes - it brought just a small increase in score. With claims, abstract and description this approach totally failed. Probably because it ignored frequency within a single document: while codes never repeated within same patent and words in titles were rarely repeating, this was no longer the case with large texts.</p>",
  "messages": [
    {
      "id": 2936037,
      "postDate": "2024-07-25T18:09:32Z",
      "content": "<p>The idea of my solution was very simple: using document frequencies to select the most important tokens and then join top-25 most important tokens with \"OR\" operator.</p>\n<p>Calculating document frequencies for CPC codes was easy and computationally not demanding, so I've made it inline in the submission:</p>\n<pre><code>code_freq = pl.read.explode.value.sort('count', descending=True)  \n\ncode_freq = code_freq.rows\n</code></pre>\n<p>Titles were a little bit more heavy, but still not worth a separate notebook:</p>\n<pre><code>vectorizer = (max_df=, min_df=, binary=True)\n\ntitles_df = pl()\\\n        ()\\\n        ()\n\nword_title_freq = vectorizer(titles_df)\n\nword_title_freq = (\n    (\n        vectorizer(), \n        np(np(word_title_freq(axis=)))\n    )\n)\n</code></pre>\n<p>Having these global document frequencies for CPC codes and words in titles, I just calculated local document frequencies for patents who were neighbors and calculated importance of a token as <code>local_document_frequency / global_document_frequency</code>:</p>\n<pre><code> = [code for codes in row[] for code in codes ]\n = Counter(flat_list_codes)\n = [(f, count / code_freq.get(elem, )) for (elem, count) in codes_counter.items()]\n\n = [title for title in row[] if title]\n = [token.text for title in titles for token in analyzer(title)]\n = Counter(flat_list_title_words)\n = [(f, count / word_title_freq.get(elem, )) for (elem, count) in title_words_counter.items()]\n</code></pre>\n<p>Then just sorted all the tokens by their importance and combined via OR:</p>\n<pre><code>selected_operands = sorted(\n            weighted_codes + weighted_title_words,\n            =lambda x: x[1], \n            =)[:25]\n\nreturn .join(selected_operands)\n</code></pre>\n<p>I've tried experimenting with same approach for pairs of codes - it brought just a small increase in score. With claims, abstract and description this approach totally failed. Probably because it ignored frequency within a single document: while codes never repeated within same patent and words in titles were rarely repeating, this was no longer the case with large texts.</p>",
      "rawMarkdown": "The idea of my solution was very simple: using document frequencies to select the most important tokens and then join top-25 most important tokens with \"OR\" operator.\n\nCalculating document frequencies for CPC codes was easy and computationally not demanding, so I've made it inline in the submission:\n```\ncode_freq = pl.read_parquet('/kaggle/input/uspto-explainable-ai/patent_metadata.parquet')['cpc_codes'].explode().value_counts().sort('count', descending=True)  \n\ncode_freq = code_freq.rows_by_key(key=[\"cpc_codes\"], unique=True)\n```\n\nTitles were a little bit more heavy, but still not worth a separate notebook:\n```\nvectorizer = CountVectorizer(max_df=10000, min_df=10, binary=True)\n\ntitles_df = pl.scan_parquet('/kaggle/input/uspto-explainable-ai/patent_data/*')\\\n        .select(['publication_number', 'title'])\\\n        .collect()\n\nword_title_freq = vectorizer.fit_transform(titles_df['title'])\n\nword_title_freq = dict(\n    zip(\n        vectorizer.get_feature_names_out(), \n        np.squeeze(np.asarray(word_title_freq.sum(axis=0)))\n    )\n)\n```\n\nHaving these global document frequencies for CPC codes and words in titles, I just calculated local document frequencies for patents who were neighbors and calculated importance of a token as `local_document_frequency / global_document_frequency`:\n```\nflat_list_codes = [code for codes in row['cpc_codes'] for code in codes ]\ncodes_counter = Counter(flat_list_codes)\nweighted_codes = [(f'cpc:{elem}', count / code_freq.get(elem, 1)) for (elem, count) in codes_counter.items()]\n\ntitles = [title for title in row['title'] if title]\nflat_list_title_words = [token.text for title in titles for token in analyzer(title)]\ntitle_words_counter = Counter(flat_list_title_words)\nweighted_title_words = [(f'ti:{elem}', count / word_title_freq.get(elem, 10000)) for (elem, count) in title_words_counter.items()]\n```\n\nThen just sorted all the tokens by their importance and combined via OR:\n```\nselected_operands = sorted(\n            weighted_codes + weighted_title_words,\n            key=lambda x: x[1], \n            reverse=True)[:25]\n\nreturn ' OR '.join(selected_operands)\n```\n\nI've tried experimenting with same approach for pairs of codes - it brought just a small increase in score. With claims, abstract and description this approach totally failed. Probably because it ignored frequency within a single document: while codes never repeated within same patent and words in titles were rarely repeating, this was no longer the case with large texts.",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2936037": "The idea of my solution was very simple: using document frequencies to select the most important tokens and then join top-25 most important tokens with \"OR\" operator.\n\nCalculating document frequencies for CPC codes was easy and computationally not demanding, so I've made it inline in the submission:\n```\ncode_freq = pl.read_parquet('/kaggle/input/uspto-explainable-ai/patent_metadata.parquet')['cpc_codes'].explode().value_counts().sort('count', descending=True)  \n\ncode_freq = code_freq.rows_by_key(key=[\"cpc_codes\"], unique=True)\n```\n\nTitles were a little bit more heavy, but still not worth a separate notebook:\n```\nvectorizer = CountVectorizer(max_df=10000, min_df=10, binary=True)\n\ntitles_df = pl.scan_parquet('/kaggle/input/uspto-explainable-ai/patent_data/*')\\\n        .select(['publication_number', 'title'])\\\n        .collect()\n\nword_title_freq = vectorizer.fit_transform(titles_df['title'])\n\nword_title_freq = dict(\n    zip(\n        vectorizer.get_feature_names_out(), \n        np.squeeze(np.asarray(word_title_freq.sum(axis=0)))\n    )\n)\n```\n\nHaving these global document frequencies for CPC codes and words in titles, I just calculated local document frequencies for patents who were neighbors and calculated importance of a token as `local_document_frequency / global_document_frequency`:\n```\nflat_list_codes = [code for codes in row['cpc_codes'] for code in codes ]\ncodes_counter = Counter(flat_list_codes)\nweighted_codes = [(f'cpc:{elem}', count / code_freq.get(elem, 1)) for (elem, count) in codes_counter.items()]\n\ntitles = [title for title in row['title'] if title]\nflat_list_title_words = [token.text for title in titles for token in analyzer(title)]\ntitle_words_counter = Counter(flat_list_title_words)\nweighted_title_words = [(f'ti:{elem}', count / word_title_freq.get(elem, 10000)) for (elem, count) in title_words_counter.items()]\n```\n\nThen just sorted all the tokens by their importance and combined via OR:\n```\nselected_operands = sorted(\n            weighted_codes + weighted_title_words,\n            key=lambda x: x[1], \n            reverse=True)[:25]\n\nreturn ' OR '.join(selected_operands)\n```\n\nI've tried experimenting with same approach for pairs of codes - it brought just a small increase in score. With claims, abstract and description this approach totally failed. Probably because it ignored frequency within a single document: while codes never repeated within same patent and words in titles were rarely repeating, this was no longer the case with large texts."
  }
}