{
  "id": 500339,
  "title": "Did someone make a working title-related submission yet?",
  "url": "/competitions/uspto-explainable-ai/discussion/500339",
  "author_name": "",
  "post_date": "2024-05-05T07:15:44.646949300Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It seems, this is a very unique and interesting challenge 😃. However, I face a lot of submission errors. Basically, I was just able to submit a single cpc-related query yet.</p>\n<p>I tried to utilize the titles now. In a first step, I load the relevant titles from the single parquet files:</p>\n<pre><code>PATENT_DATA = \npatent_details = (os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\n patent_detail  tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=[,])\n    tmp = tmp.loc[tmp[].isin(allunique_pubnums)].reset_index(drop=)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=).reset_index(drop=)`\n</code></pre>\n<p>After I build the queries, I use the most frequent word of the titles as very sample baseline:</p>\n<pre><code> ():\n     .join([+x.replace(,)  x  keywords]) \n\nTOP_N = \nquery_df = pd.DataFrame()\nqueries_ti = []\n\n i, row  test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df[].isin(row[rel_columns].values)][].copy()\n    tmp_text = .join(single_sample.values).replace(,).replace(,)\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text  token  custom_analyzer(.join((word_counts.keys())))]\n    tokens = ((tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n</code></pre>\n<p>The queries look like this:</p>\n<table>\n<thead>\n<tr>\n<th>publication_number</th>\n<th>query</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>US-2017082634-A1</td>\n<td>ti:mass</td>\n</tr>\n<tr>\n<td>US-2017180470-A1</td>\n<td>ti:method</td>\n</tr>\n<tr>\n<td>US-2018029544-A1</td>\n<td>ti:module</td>\n</tr>\n</tbody>\n</table>\n<p>However, I can succesfully save it but I get a \"Notebook Threw Exception\" error when trying to submit it. Anyone else faced this error? I think the error must be in one of the both code blocks above but I don´t know where. It works locally and also with the provided test samples.</p>",
  "messages": [
    {
      "id": "2794257",
      "postDate": "05/05/2024 07:15:44",
      "content": "<p>It seems, this is a very unique and interesting challenge 😃. However, I face a lot of submission errors. Basically, I was just able to submit a single cpc-related query yet.</p>\n<p>I tried to utilize the titles now. In a first step, I load the relevant titles from the single parquet files:</p>\n<pre><code>PATENT_DATA = \npatent_details = (os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\n patent_detail  tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=[,])\n    tmp = tmp.loc[tmp[].isin(allunique_pubnums)].reset_index(drop=)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=).reset_index(drop=)`\n</code></pre>\n<p>After I build the queries, I use the most frequent word of the titles as very sample baseline:</p>\n<pre><code> ():\n     .join([+x.replace(,)  x  keywords]) \n\nTOP_N = \nquery_df = pd.DataFrame()\nqueries_ti = []\n\n i, row  test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df[].isin(row[rel_columns].values)][].copy()\n    tmp_text = .join(single_sample.values).replace(,).replace(,)\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text  token  custom_analyzer(.join((word_counts.keys())))]\n    tokens = ((tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n</code></pre>\n<p>The queries look like this:</p>\n<table>\n<thead>\n<tr>\n<th>publication_number</th>\n<th>query</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>US-2017082634-A1</td>\n<td>ti:mass</td>\n</tr>\n<tr>\n<td>US-2017180470-A1</td>\n<td>ti:method</td>\n</tr>\n<tr>\n<td>US-2018029544-A1</td>\n<td>ti:module</td>\n</tr>\n</tbody>\n</table>\n<p>However, I can succesfully save it but I get a \"Notebook Threw Exception\" error when trying to submit it. Anyone else faced this error? I think the error must be in one of the both code blocks above but I don´t know where. It works locally and also with the provided test samples.</p>",
      "rawMarkdown": "It seems, this is a very unique and interesting challenge 😃. However, I face a lot of submission errors. Basically, I was just able to submit a single cpc-related query yet.\n\nI tried to utilize the titles now. In a first step, I load the relevant titles from the single parquet files:\n\n```python\nPATENT_DATA = '/kaggle/input/uspto-explainable-ai/patent_data'\npatent_details = sorted(os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\nfor patent_detail in tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=['publication_number','title'])\n    tmp = tmp.loc[tmp['publication_number'].isin(allunique_pubnums)].reset_index(drop=True)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=0).reset_index(drop=True)`\n```\n\nAfter I build the queries, I use the most frequent word of the titles as very sample baseline:\n\n```python\ndef build_query(keywords):\n    return \" AND \".join(['ti:'+x.replace(\"\\\\\",\"\") for x in keywords]) # \n\nTOP_N = 1\nquery_df = pd.DataFrame()\nqueries_ti = []\n\nfor i, row in test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df['publication_number'].isin(row[rel_columns].values)]['title'].copy()\n    tmp_text = \" \".join(single_sample.values).replace(\",\",\"\").replace(\".\",\"\")\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text for token in custom_analyzer(\" \".join(list(word_counts.keys())))]\n    tokens = list(set(tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n```\n\nThe queries look like this:\n| publication_number | query |\n| --- | --- |\n| US-2017082634-A1 | ti:mass |\n| US-2017180470-A1 | ti:method |\n| US-2018029544-A1 | ti:module |\n\nHowever, I can succesfully save it but I get a \"Notebook Threw Exception\" error when trying to submit it. Anyone else faced this error? I think the error must be in one of the both code blocks above but I don´t know where. It works locally and also with the provided test samples.",
      "votes": null
    },
    {
      "id": "2794623",
      "postDate": "05/05/2024 12:12:30",
      "content": "<p><a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a> informed i.e working fine with multiple keywords combinations go through other discussions once.</p>",
      "rawMarkdown": "devinanzelmo informed i.e working fine with multiple keywords combinations go through other discussions once.",
      "votes": null
    },
    {
      "id": "2794655",
      "postDate": "05/05/2024 12:36:06",
      "content": "<p>Thank you very much. However, I can confirm that the \"bug\" is in this part:</p>\n<pre><code> i, row  test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df[].isin(row[rel_columns].values)][].copy()\n    tmp_text = .join(single_sample.values).replace(,).replace(,)\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text  token  custom_analyzer(.join((word_counts.keys())))]\n    tokens = ((tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n</code></pre>\n<p>Step by step:</p>\n<ul>\n<li>Initially I read all relevant patent metadata from the parquet files to get the titles and combine it to a dataframe:</li>\n</ul>\n<pre><code>patent_details = (os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\n patent_detail  tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=[,])\n    tmp = tmp.loc[tmp[].isin(allunique_pubnums)].reset_index(drop=)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=).reset_index(drop=)\n</code></pre>\n<ul>\n<li>After I iterate over test_df and get the titles for each of the targets</li>\n<li>I count the most frequent words, remove stopwords and build a simple query like \"ti:MOST_FREQUENT_WORD\"</li>\n</ul>\n<p>If I integrate a try/except block for this part, code will run through but looking at the score, the exception is triggered for each row. I don´t know why it doesn´t work on the hidden test set/notebook.</p>",
      "rawMarkdown": "Thank you very much. However, I can confirm that the \"bug\" is in this part:\n\n```python\nfor i, row in test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df['publication_number'].isin(row[rel_columns].values)]['title'].copy()\n    tmp_text = \" \".join(single_sample.values).replace(\",\",\"\").replace(\".\",\"\")\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text for token in custom_analyzer(\" \".join(list(word_counts.keys())))]\n    tokens = list(set(tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n```\n\nStep by step:\n- Initially I read all relevant patent metadata from the parquet files to get the titles and combine it to a dataframe:\n\n```python\npatent_details = sorted(os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\nfor patent_detail in tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=['publication_number','title'])\n    tmp = tmp.loc[tmp['publication_number'].isin(allunique_pubnums)].reset_index(drop=True)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=0).reset_index(drop=True)\n```\n\n- After I iterate over test_df and get the titles for each of the targets\n- I count the most frequent words, remove stopwords and build a simple query like \"ti:MOST_FREQUENT_WORD\"\n\nIf I integrate a try/except block for this part, code will run through but looking at the score, the exception is triggered for each row. I don´t know why it doesn´t work on the hidden test set/notebook.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2794623,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "05/05/2024 12:12:30",
      "content": "<p><a href=\"https://www.kaggle.com/devinanzelmo\" target=\"_blank\">@devinanzelmo</a> informed i.e working fine with multiple keywords combinations go through other discussions once.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2794655,
          "author_name": "benbla",
          "author_url": "",
          "post_date": "05/05/2024 12:36:06",
          "content": "<p>Thank you very much. However, I can confirm that the \"bug\" is in this part:</p>\n<pre><code> i, row  test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df[].isin(row[rel_columns].values)][].copy()\n    tmp_text = .join(single_sample.values).replace(,).replace(,)\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text  token  custom_analyzer(.join((word_counts.keys())))]\n    tokens = ((tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n</code></pre>\n<p>Step by step:</p>\n<ul>\n<li>Initially I read all relevant patent metadata from the parquet files to get the titles and combine it to a dataframe:</li>\n</ul>\n<pre><code>patent_details = (os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\n patent_detail  tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=[,])\n    tmp = tmp.loc[tmp[].isin(allunique_pubnums)].reset_index(drop=)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=).reset_index(drop=)\n</code></pre>\n<ul>\n<li>After I iterate over test_df and get the titles for each of the targets</li>\n<li>I count the most frequent words, remove stopwords and build a simple query like \"ti:MOST_FREQUENT_WORD\"</li>\n</ul>\n<p>If I integrate a try/except block for this part, code will run through but looking at the score, the exception is triggered for each row. I don´t know why it doesn´t work on the hidden test set/notebook.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2794257": "It seems, this is a very unique and interesting challenge 😃. However, I face a lot of submission errors. Basically, I was just able to submit a single cpc-related query yet.\n\nI tried to utilize the titles now. In a first step, I load the relevant titles from the single parquet files:\n\n```python\nPATENT_DATA = '/kaggle/input/uspto-explainable-ai/patent_data'\npatent_details = sorted(os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\nfor patent_detail in tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=['publication_number','title'])\n    tmp = tmp.loc[tmp['publication_number'].isin(allunique_pubnums)].reset_index(drop=True)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=0).reset_index(drop=True)`\n```\n\nAfter I build the queries, I use the most frequent word of the titles as very sample baseline:\n\n```python\ndef build_query(keywords):\n    return \" AND \".join(['ti:'+x.replace(\"\\\\\",\"\") for x in keywords]) # \n\nTOP_N = 1\nquery_df = pd.DataFrame()\nqueries_ti = []\n\nfor i, row in test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df['publication_number'].isin(row[rel_columns].values)]['title'].copy()\n    tmp_text = \" \".join(single_sample.values).replace(\",\",\"\").replace(\".\",\"\")\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text for token in custom_analyzer(\" \".join(list(word_counts.keys())))]\n    tokens = list(set(tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n```\n\nThe queries look like this:\n| publication_number | query |\n| --- | --- |\n| US-2017082634-A1 | ti:mass |\n| US-2017180470-A1 | ti:method |\n| US-2018029544-A1 | ti:module |\n\nHowever, I can succesfully save it but I get a \"Notebook Threw Exception\" error when trying to submit it. Anyone else faced this error? I think the error must be in one of the both code blocks above but I don´t know where. It works locally and also with the provided test samples.",
    "2794623": "devinanzelmo informed i.e working fine with multiple keywords combinations go through other discussions once.",
    "2794655": "Thank you very much. However, I can confirm that the \"bug\" is in this part:\n\n```python\nfor i, row in test_df.iterrows():\n    single_sample = patent_detail_df.loc[patent_detail_df['publication_number'].isin(row[rel_columns].values)]['title'].copy()\n    tmp_text = \" \".join(single_sample.values).replace(\",\",\"\").replace(\".\",\"\")\n    word_counts = pd.Series(tmp_text.lower().split()).value_counts().to_dict()\n    tokens = [token.text for token in custom_analyzer(\" \".join(list(word_counts.keys())))]\n    tokens = list(set(tokens))[:TOP_N]\n    queries_ti.append(build_query(tokens))\n```\n\nStep by step:\n- Initially I read all relevant patent metadata from the parquet files to get the titles and combine it to a dataframe:\n\n```python\npatent_details = sorted(os.listdir(PATENT_DATA))\npatent_detail_df = pd.DataFrame()\nfor patent_detail in tqdm(patent_details):\n    tmp = pd.read_parquet(os.path.join(PATENT_DATA, patent_detail), columns=['publication_number','title'])\n    tmp = tmp.loc[tmp['publication_number'].isin(allunique_pubnums)].reset_index(drop=True)\n    patent_detail_df = pd.concat([patent_detail_df, tmp], axis=0).reset_index(drop=True)\n```\n\n- After I iterate over test_df and get the titles for each of the targets\n- I count the most frequent words, remove stopwords and build a simple query like \"ti:MOST_FREQUENT_WORD\"\n\nIf I integrate a try/except block for this part, code will run through but looking at the score, the exception is triggered for each row. I don´t know why it doesn´t work on the hidden test set/notebook."
  },
  "source": "meta"
}