{
  "id": 521513,
  "title": "Unable to submit due to error",
  "url": "/competitions/uspto-explainable-ai/discussion/521513",
  "author_name": "",
  "post_date": "2024-07-21T10:01:46.409148Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm trying to submit using the code below, but I can't submit due to an error. Somebody please help me.</p>\n<blockquote>\n  <p>import torch<br>\n  from transformers import BertModel, BertTokenizer, BertConfig<br>\n  import numpy as np<br>\n  import polars as pl<br>\n  from tqdm import tqdm<br>\n  import whoosh_utils<br>\n  from pathlib import Path<br>\n  from dataclasses import dataclass<br>\n  from typing import Any, List<br>\n  import random<br>\n  import signal<br>\n  import time<br>\n  import datetime<br>\n  import pickle<br>\n  import copy<br>\n  model_path = '/kaggle/input/berdmodel-2'<br>\n  tokenizer_path = '/kaggle/input/berdmodel-2'<br>\n  config = BertConfig.from_pretrained(model_path, output_attentions=True)<br>\n  model = BertModel.from_pretrained(model_path, config=config)<br>\n  tokenizer = BertTokenizer.from_pretrained(tokenizer_path)<br>\n  device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")<br>\n  model.to(device)<br>\n  def identity(x):<br>\n      return x<br>\n  with open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:<br>\n      cpc_cv_tfidf = pickle.load(f)<br>\n  def select_top_k_columns(X: Any, k: int) -&gt; tuple[Any, np.ndarray]:<br>\n      # 行方向の和を計算<br>\n      row_sums = X.sum(axis=0)<br>\n      # 和が大きい上位k個の列のインデックスを取得<br>\n      top_k_indices = np.argsort(-row_sums.A1)[:k]<br>\n      # 上位k個の列を選択<br>\n      X_top = X[:, top_k_indices]<br>\n      return X_top, top_k_indices<br>\n  def extract_important_tokens(text, model, tokenizer, num_tokens=5):<br>\n      inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)<br>\n      with torch.no_grad():<br>\n          outputs = model(**inputs)<br>\n          last_hidden_state = outputs.last_hidden_state<br>\n          attention_weights = outputs.attentions[-1]  # 最後の注意層を取得<br>\n      attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()<br>\n      tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())<br>\n      filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]<br>\n      sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)<br>\n      unique_important_tokens = []<br>\n      seen_tokens = set()<br>\n      for token in filtered_tokens:<br>\n          if len(token[0]) &gt; 1 and token[0] not in seen_tokens:<br>\n              unique_important_tokens.append(token[0])<br>\n              seen_tokens.add(token[0])<br>\n          if len(unique_important_tokens) &gt;= num_tokens:<br>\n              break<br>\n      return unique_important_tokens<br>\n  <a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\n  class Word:<br>\n      category: str<br>\n      content: str<br>\n      def to_str(self):<br>\n          return f\"{self.category}:{self.content}\"</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\nclass State:<br>\n    words: List[Word]</p>\n<pre><code> ():\n    self.use = np.random.binomial(, , (self.words))\n\n ():\n    words = [word.to_str()  word, use  (self.words, self.use)  use]\n     (words) &gt; :  \n        words = words[:]\n     .join(words)\n\n ():\n    idx = np.random.choice((self.words))\n    self.use[idx] =  - self.use[idx]\n     self\n</code></pre>\n<p>class USPTOProblem(Annealer):<br>\n    def <strong>init</strong>(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):<br>\n        super(USPTOProblem, self).<strong>init</strong>(init_state)<br>\n        self.qp = qp<br>\n        self.searcher = searcher<br>\n        self.target = target<br>\n        self.Tmax = tmax<br>\n        self.Tmin = tmin<br>\n        self.steps = steps<br>\n        self.max_time = max_time<br>\n        self.copy_strategy = copy_strategy</p>\n<pre><code>def move():\n    ..move_1()\n\ndef energy():\n    query = ..to_query()\n    cand = whoosh_utils.execute_query(query, .qp, .searcher)\n    ap50_score = ap50(cand, .target)\n    return -ap50_score\n</code></pre>\n<p>def ap50(preds: list[str], labels: list[str]) -&gt; float:<br>\n    precisions = []<br>\n    n_found = 0<br>\n    for e, i in enumerate(preds):<br>\n        if i in labels:<br>\n            n_found += 1<br>\n        precisions.append(n_found / (e + 1))<br>\n    return sum(precisions) / 50</p>\n<p>test_idx = whoosh_utils.load_index(\"./test_index\")<br>\nsearcher = whoosh_utils.get_searcher(test_idx)<br>\nqp = whoosh_utils.get_query_parser()</p>\n<p>scores = []<br>\nresults = []</p>\n<p>for i in tqdm(range(len(test))):<br>\n    target = test[i].to_numpy().flatten()[1:].tolist()<br>\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))</p>\n<pre><code> len(meta_i) == 0:\n    results.append({: test[i, ], : })\n    (, i)\n    continue\n\ntitles = meta_i.get_column().fill_null().to_list()\nabstracts = meta_i.get_column().fill_null().to_list()\ncpc_lists = meta_i.get_column().to_list()\ncpcs = [cpc  sublist  cpc_lists  cpc  sublist]  # リスト内のリストを平坦化\n\n\ncpc_mat = cpc_cv_tfidf.transform(cpcs)\nX_cpc, cpc_idx = select_top_k_columns(cpc_mat, =3)\ntopk_cpc = cpc_cv_tfidf.get_feature_names_out()[cpc_idx]\ntopk_cpc = [Word(=, =x)  x  topk_cpc]\n\ntopk_titles = []\ntopk_abstracts = []\n title  titles:\n    topk_titles.extend(extract_important_tokens(title, model, tokenizer, =10))\n abstract  abstracts:\n    topk_abstracts.extend(extract_important_tokens(abstract, model, tokenizer, =30))\n\n\ntopk_titles = list((topk_titles))[:30]\ntopk_abstracts = list((topk_abstracts))[:30]\ntopk_words = [Word(=, =x)  x  topk_titles]\ntopk_words += [Word(=, =x)  x  topk_abstracts]\nwords = topk_words + topk_cpc\nstate = State(=words)\nproblem = USPTOProblem(qp, searcher, target, state, =1000, =10)\nsolution, score = problem.anneal()\n(f, -score)\nscores.append(-score)\nresults.append({: test[i, ], : solution.to_query()})\n</code></pre>\n<p>print(\"Average Score:\", sum(scores) / len(scores))</p>",
  "messages": [
    {
      "id": "2930693",
      "postDate": "07/21/2024 10:01:46",
      "content": "<p>I'm trying to submit using the code below, but I can't submit due to an error. Somebody please help me.</p>\n<blockquote>\n  <p>import torch<br>\n  from transformers import BertModel, BertTokenizer, BertConfig<br>\n  import numpy as np<br>\n  import polars as pl<br>\n  from tqdm import tqdm<br>\n  import whoosh_utils<br>\n  from pathlib import Path<br>\n  from dataclasses import dataclass<br>\n  from typing import Any, List<br>\n  import random<br>\n  import signal<br>\n  import time<br>\n  import datetime<br>\n  import pickle<br>\n  import copy<br>\n  model_path = '/kaggle/input/berdmodel-2'<br>\n  tokenizer_path = '/kaggle/input/berdmodel-2'<br>\n  config = BertConfig.from_pretrained(model_path, output_attentions=True)<br>\n  model = BertModel.from_pretrained(model_path, config=config)<br>\n  tokenizer = BertTokenizer.from_pretrained(tokenizer_path)<br>\n  device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")<br>\n  model.to(device)<br>\n  def identity(x):<br>\n      return x<br>\n  with open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:<br>\n      cpc_cv_tfidf = pickle.load(f)<br>\n  def select_top_k_columns(X: Any, k: int) -&gt; tuple[Any, np.ndarray]:<br>\n      # 行方向の和を計算<br>\n      row_sums = X.sum(axis=0)<br>\n      # 和が大きい上位k個の列のインデックスを取得<br>\n      top_k_indices = np.argsort(-row_sums.A1)[:k]<br>\n      # 上位k個の列を選択<br>\n      X_top = X[:, top_k_indices]<br>\n      return X_top, top_k_indices<br>\n  def extract_important_tokens(text, model, tokenizer, num_tokens=5):<br>\n      inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)<br>\n      with torch.no_grad():<br>\n          outputs = model(**inputs)<br>\n          last_hidden_state = outputs.last_hidden_state<br>\n          attention_weights = outputs.attentions[-1]  # 最後の注意層を取得<br>\n      attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()<br>\n      tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())<br>\n      filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]<br>\n      sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)<br>\n      unique_important_tokens = []<br>\n      seen_tokens = set()<br>\n      for token in filtered_tokens:<br>\n          if len(token[0]) &gt; 1 and token[0] not in seen_tokens:<br>\n              unique_important_tokens.append(token[0])<br>\n              seen_tokens.add(token[0])<br>\n          if len(unique_important_tokens) &gt;= num_tokens:<br>\n              break<br>\n      return unique_important_tokens<br>\n  <a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\n  class Word:<br>\n      category: str<br>\n      content: str<br>\n      def to_str(self):<br>\n          return f\"{self.category}:{self.content}\"</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\nclass State:<br>\n    words: List[Word]</p>\n<pre><code> ():\n    self.use = np.random.binomial(, , (self.words))\n\n ():\n    words = [word.to_str()  word, use  (self.words, self.use)  use]\n     (words) &gt; :  \n        words = words[:]\n     .join(words)\n\n ():\n    idx = np.random.choice((self.words))\n    self.use[idx] =  - self.use[idx]\n     self\n</code></pre>\n<p>class USPTOProblem(Annealer):<br>\n    def <strong>init</strong>(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):<br>\n        super(USPTOProblem, self).<strong>init</strong>(init_state)<br>\n        self.qp = qp<br>\n        self.searcher = searcher<br>\n        self.target = target<br>\n        self.Tmax = tmax<br>\n        self.Tmin = tmin<br>\n        self.steps = steps<br>\n        self.max_time = max_time<br>\n        self.copy_strategy = copy_strategy</p>\n<pre><code>def move():\n    ..move_1()\n\ndef energy():\n    query = ..to_query()\n    cand = whoosh_utils.execute_query(query, .qp, .searcher)\n    ap50_score = ap50(cand, .target)\n    return -ap50_score\n</code></pre>\n<p>def ap50(preds: list[str], labels: list[str]) -&gt; float:<br>\n    precisions = []<br>\n    n_found = 0<br>\n    for e, i in enumerate(preds):<br>\n        if i in labels:<br>\n            n_found += 1<br>\n        precisions.append(n_found / (e + 1))<br>\n    return sum(precisions) / 50</p>\n<p>test_idx = whoosh_utils.load_index(\"./test_index\")<br>\nsearcher = whoosh_utils.get_searcher(test_idx)<br>\nqp = whoosh_utils.get_query_parser()</p>\n<p>scores = []<br>\nresults = []</p>\n<p>for i in tqdm(range(len(test))):<br>\n    target = test[i].to_numpy().flatten()[1:].tolist()<br>\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))</p>\n<pre><code> len(meta_i) == 0:\n    results.append({: test[i, ], : })\n    (, i)\n    continue\n\ntitles = meta_i.get_column().fill_null().to_list()\nabstracts = meta_i.get_column().fill_null().to_list()\ncpc_lists = meta_i.get_column().to_list()\ncpcs = [cpc  sublist  cpc_lists  cpc  sublist]  # リスト内のリストを平坦化\n\n\ncpc_mat = cpc_cv_tfidf.transform(cpcs)\nX_cpc, cpc_idx = select_top_k_columns(cpc_mat, =3)\ntopk_cpc = cpc_cv_tfidf.get_feature_names_out()[cpc_idx]\ntopk_cpc = [Word(=, =x)  x  topk_cpc]\n\ntopk_titles = []\ntopk_abstracts = []\n title  titles:\n    topk_titles.extend(extract_important_tokens(title, model, tokenizer, =10))\n abstract  abstracts:\n    topk_abstracts.extend(extract_important_tokens(abstract, model, tokenizer, =30))\n\n\ntopk_titles = list((topk_titles))[:30]\ntopk_abstracts = list((topk_abstracts))[:30]\ntopk_words = [Word(=, =x)  x  topk_titles]\ntopk_words += [Word(=, =x)  x  topk_abstracts]\nwords = topk_words + topk_cpc\nstate = State(=words)\nproblem = USPTOProblem(qp, searcher, target, state, =1000, =10)\nsolution, score = problem.anneal()\n(f, -score)\nscores.append(-score)\nresults.append({: test[i, ], : solution.to_query()})\n</code></pre>\n<p>print(\"Average Score:\", sum(scores) / len(scores))</p>",
      "rawMarkdown": "I'm trying to submit using the code below, but I can't submit due to an error. Somebody please help me.\n\n\n>import torch\nfrom transformers import BertModel, BertTokenizer, BertConfig\nimport numpy as np\nimport polars as pl\nfrom tqdm import tqdm\nimport whoosh_utils\nfrom pathlib import Path\nfrom dataclasses import dataclass\nfrom typing import Any, List\nimport random\nimport signal\nimport time\nimport datetime\nimport pickle\nimport copy\nmodel_path = '/kaggle/input/berdmodel-2'\ntokenizer_path = '/kaggle/input/berdmodel-2'\nconfig = BertConfig.from_pretrained(model_path, output_attentions=True)\nmodel = BertModel.from_pretrained(model_path, config=config)\ntokenizer = BertTokenizer.from_pretrained(tokenizer_path)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel.to(device)\ndef identity(x):\n    return x\nwith open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:\n    cpc_cv_tfidf = pickle.load(f)\ndef select_top_k_columns(X: Any, k: int) -> tuple[Any, np.ndarray]:\n    # 行方向の和を計算\n    row_sums = X.sum(axis=0)\n    # 和が大きい上位k個の列のインデックスを取得\n    top_k_indices = np.argsort(-row_sums.A1)[:k]\n    # 上位k個の列を選択\n    X_top = X[:, top_k_indices]\n    return X_top, top_k_indices\ndef extract_important_tokens(text, model, tokenizer, num_tokens=5):\n    inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)\n    with torch.no_grad():\n        outputs = model(**inputs)\n        last_hidden_state = outputs.last_hidden_state\n        attention_weights = outputs.attentions[-1]  # 最後の注意層を取得\n    attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()\n    tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())\n    filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]\n    sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)\n    unique_important_tokens = []\n    seen_tokens = set()\n    for token in filtered_tokens:\n        if len(token[0]) > 1 and token[0] not in seen_tokens:\n            unique_important_tokens.append(token[0])\n            seen_tokens.add(token[0])\n        if len(unique_important_tokens) >= num_tokens:\n            break\n    return unique_important_tokens\n@dataclass\nclass Word:\n    category: str\n    content: str\n    def to_str(self):\n        return f\"{self.category}:{self.content}\"\n\n@dataclass\nclass State:\n    words: List[Word]\n\n    def __post_init__(self):\n        self.use = np.random.binomial(1, 0.5, len(self.words))\n\n    def to_query(self):\n        words = [word.to_str() for word, use in zip(self.words, self.use) if use]\n        if len(words) > 10000:  # クエリの長さを制限\n            words = words[:10000]\n        return \" OR \".join(words)\n\n    def move_1(self):\n        idx = np.random.choice(len(self.words))\n        self.use[idx] = 1 - self.use[idx]\n        return self\n\nclass USPTOProblem(Annealer):\n    def __init__(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):\n        super(USPTOProblem, self).__init__(init_state)\n        self.qp = qp\n        self.searcher = searcher\n        self.target = target\n        self.Tmax = tmax\n        self.Tmin = tmin\n        self.steps = steps\n        self.max_time = max_time\n        self.copy_strategy = copy_strategy\n\n    def move(self):\n        self.state.move_1()\n\n    def energy(self):\n        query = self.state.to_query()\n        cand = whoosh_utils.execute_query(query, self.qp, self.searcher)\n        ap50_score = ap50(cand, self.target)\n        return -ap50_score\n\ndef ap50(preds: list[str], labels: list[str]) -> float:\n    precisions = []\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n        precisions.append(n_found / (e + 1))\n    return sum(precisions) / 50\n\ntest_idx = whoosh_utils.load_index(\"./test_index\")\nsearcher = whoosh_utils.get_searcher(test_idx)\nqp = whoosh_utils.get_query_parser()\n\nscores = []\nresults = []\n\nfor i in tqdm(range(len(test))):\n    target = test[i].to_numpy().flatten()[1:].tolist()\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))\n\n    if len(meta_i) == 0:\n        results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": \"ti:device\"})\n        print(\"\\t Append Dummy\", i)\n        continue\n\n    titles = meta_i.get_column(\"title\").fill_null(\"\").to_list()\n    abstracts = meta_i.get_column(\"abstract\").fill_null(\"\").to_list()\n    cpc_lists = meta_i.get_column(\"cpc\").to_list()\n    cpcs = [cpc for sublist in cpc_lists for cpc in sublist]  # リスト内のリストを平坦化\n\n    # TF-IDF matrix for CPC codes\n    cpc_mat = cpc_cv_tfidf.transform(cpcs)\n    X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=3)\n    topk_cpc = cpc_cv_tfidf.get_feature_names_out()[cpc_idx]\n    topk_cpc = [Word(category=\"cpc\", content=x) for x in topk_cpc]\n    # Important topk words from titles and abstracts\n    topk_titles = []\n    topk_abstracts = []\n    for title in titles:\n        topk_titles.extend(extract_important_tokens(title, model, tokenizer, num_tokens=10))\n    for abstract in abstracts:\n        topk_abstracts.extend(extract_important_tokens(abstract, model, tokenizer, num_tokens=30))\n\n    # 一意にする\n    topk_titles = list(set(topk_titles))[:30]\n    topk_abstracts = list(set(topk_abstracts))[:30]\n    topk_words = [Word(category=\"ti\", content=x) for x in topk_titles]\n    topk_words += [Word(category=\"ab\", content=x) for x in topk_abstracts]\n    words = topk_words + topk_cpc\n    state = State(words=words)\n    problem = USPTOProblem(qp, searcher, target, state, steps=1000, max_time=10)\n    solution, score = problem.anneal()\n    print(f\"\\t Problem Number {i} Score:\", -score)\n    scores.append(-score)\n    results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": solution.to_query()})\nprint(\"Average Score:\", sum(scores) / len(scores))",
      "votes": null
    },
    {
      "id": "2931198",
      "postDate": "07/21/2024 18:14:17",
      "content": "<p>It seems that you are experiencing submission errors, likely due to issues in generating the submission file. Based on the provided code, here are the potential issues and the suggested fix:</p>\n<p>Potential Issues:<br>\nIncorrect Submission File Format:</p>\n<p>Your submission file may not meet the required format, which typically includes specific columns and data types.<br>\nResidual Files:</p>\n<p>There may be residual files in your working directory that are causing conflicts.<br>\nMissing Submission File Generation:</p>\n<p>It appears that the code to generate and save the submission file is missing.<br>\nSuggested Fix:<br>\nEnsure Correct Submission File Format:</p>\n<p>Verify that the submission file has the correct number of rows and columns, and that all required columns are present.<br>\nClean Working Directory:</p>\n<p>Remove any unnecessary files from the working directory to prevent conflicts.<br>\nGenerate and Save the Submission File:</p>\n<p>Add the code to create and save the submission file.<br>\nHere is the modified code with these changes:</p>\n<p>import torch<br>\nfrom transformers import BertModel, BertTokenizer, BertConfig<br>\nimport numpy as np<br>\nimport polars as pl<br>\nfrom tqdm import tqdm<br>\nimport whoosh_utils<br>\nfrom pathlib import Path<br>\nfrom dataclasses import dataclass<br>\nfrom typing import Any, List<br>\nimport random<br>\nimport signal<br>\nimport time<br>\nimport datetime<br>\nimport pickle<br>\nimport copy</p>\n<h1>Load model and tokenizer</h1>\n<p>model_path = '/kaggle/input/berdmodel-2'<br>\ntokenizer_path = '/kaggle/input/berdmodel-2'<br>\nconfig = BertConfig.from_pretrained(model_path, output_attentions=True)<br>\nmodel = BertModel.from_pretrained(model_path, config=config)<br>\ntokenizer = BertTokenizer.from_pretrained(tokenizer_path)<br>\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")<br>\nmodel.to(device)</p>\n<p>def identity(x):<br>\n    return x</p>\n<p>with open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:<br>\n    cpc_cv_tfidf = pickle.load(f)</p>\n<p>def select_top_k_columns(X: Any, k: int) -&gt; tuple[Any, np.ndarray]:<br>\n    row_sums = X.sum(axis=0)<br>\n    top_k_indices = np.argsort(-row_sums.A1)[:k]<br>\n    X_top = X[:, top_k_indices]<br>\n    return X_top, top_k_indices</p>\n<p>def extract_important_tokens(text, model, tokenizer, num_tokens=5):<br>\n    inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)<br>\n    with torch.no_grad():<br>\n        outputs = model(**inputs)<br>\n        last_hidden_state = outputs.last_hidden_state<br>\n        attention_weights = outputs.attentions[-1]<br>\n        attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()<br>\n    tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())<br>\n    filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]<br>\n    sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)<br>\n    unique_important_tokens = []<br>\n    seen_tokens = set()<br>\n    for token in sorted_tokens:<br>\n        if len(token[0]) &gt; 1 and token[0] not in seen_tokens:<br>\n            unique_important_tokens.append(token[0])<br>\n            seen_tokens.add(token[0])<br>\n        if len(unique_important_tokens) &gt;= num_tokens:<br>\n            break<br>\n    return unique_important_tokens</p>\n<p><a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\nclass Word:<br>\n    category: str<br>\n    content: str</p>\n<pre><code> ():\n     \n</code></pre>\n<p><a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\nclass State:<br>\n    words: List[Word]</p>\n<pre><code> ():\n    self.use = np.random.binomial(, , (self.words))\n\n ():\n    words = [word.to_str()  word, use  (self.words, self.use)  use]\n     (words) &gt; :\n        words = words[:]\n     .join(words)\n\n ():\n    idx = np.random.choice((self.words))\n    self.use[idx] =  - self.use[idx]\n     self\n</code></pre>\n<p>class USPTOProblem(Annealer):<br>\n    def <strong>init</strong>(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):<br>\n        super(USPTOProblem, self).<strong>init</strong>(init_state)<br>\n        self.qp = qp<br>\n        self.searcher = searcher<br>\n        self.target = target<br>\n        self.Tmax = tmax<br>\n        self.Tmin = tmin<br>\n        self.steps = steps<br>\n        self.max_time = max_time<br>\n        self.copy_strategy = copy_strategy</p>\n<pre><code>def move():\n    ..move_1()\n\ndef energy():\n    query = ..to_query()\n    cand = whoosh_utils.execute_query(query, .qp, .searcher)\n    ap50_score = ap50(cand, .target)\n    return -ap50_score\n</code></pre>\n<p>def ap50(preds: list[str], labels: list[str]) -&gt; float:<br>\n    precisions = []<br>\n    n_found = 0<br>\n    for e, i in enumerate(preds):<br>\n        if i in labels:<br>\n            n_found += 1<br>\n        precisions.append(n_found / (e + 1))<br>\n    return sum(precisions) / 50</p>\n<p>test_idx = whoosh_utils.load_index(\"./test_index\")<br>\nsearcher = whoosh_utils.get_searcher(test_idx)<br>\nqp = whoosh_utils.get_query_parser()</p>\n<p>scores = []<br>\nresults = []</p>\n<p>for i in tqdm(range(len(test))):<br>\n    target = test[i].to_numpy().flatten()[1:].tolist()<br>\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))</p>\n<pre><code> (meta_i) == :\n    results({: test, : })\n    (, i)\n    continue\n\ntitles = meta_i()()()\nabstracts = meta_i()()()\ncpc_lists = meta_i()()\ncpcs = \n\ncpc_mat = cpc_cv_tfidf(cpcs)\nX_cpc, cpc_idx = (cpc_mat, k=)\ntopk_cpc = cpc_cv_tfidf()\ntopk_cpc = \n\ntopk_titles = \ntopk_abstracts = \n title  titles:\n    topk_titles((title, model, tokenizer, num_tokens=))\n abstract  abstracts:\n    topk_abstracts((abstract, model, tokenizer, num_tokens=))\n\ntopk_titles = ((topk_titles))\ntopk_abstracts = ((topk_abstracts))\ntopk_words = \ntopk_words += \nwords = topk_words + topk_cpc\nstate = (words=words)\nproblem = (qp, searcher, target, state, steps=, max_time=)\nsolution, score = problem()\n\nscores(-score)\nresults({: test, : solution()})\n</code></pre>\n<p>print(\"Average Score:\", sum(scores) / len(scores))</p>\n<h1>Clean the working directory to avoid submission errors</h1>\n<p>!rm -rf /kaggle/working/*</p>\n<h1>Generate the submission file</h1>\n<p>submission = pl.DataFrame(results)</p>\n<h1>Ensure there are no missing values</h1>\n<p>submission = submission.fill_null(\"\")</p>\n<h1>Save the submission file</h1>\n<p>submission.write_csv(\"submission.csv\")</p>",
      "rawMarkdown": "It seems that you are experiencing submission errors, likely due to issues in generating the submission file. Based on the provided code, here are the potential issues and the suggested fix:\n\nPotential Issues:\nIncorrect Submission File Format:\n\nYour submission file may not meet the required format, which typically includes specific columns and data types.\nResidual Files:\n\nThere may be residual files in your working directory that are causing conflicts.\nMissing Submission File Generation:\n\nIt appears that the code to generate and save the submission file is missing.\nSuggested Fix:\nEnsure Correct Submission File Format:\n\nVerify that the submission file has the correct number of rows and columns, and that all required columns are present.\nClean Working Directory:\n\nRemove any unnecessary files from the working directory to prevent conflicts.\nGenerate and Save the Submission File:\n\nAdd the code to create and save the submission file.\nHere is the modified code with these changes:\n\nimport torch\nfrom transformers import BertModel, BertTokenizer, BertConfig\nimport numpy as np\nimport polars as pl\nfrom tqdm import tqdm\nimport whoosh_utils\nfrom pathlib import Path\nfrom dataclasses import dataclass\nfrom typing import Any, List\nimport random\nimport signal\nimport time\nimport datetime\nimport pickle\nimport copy\n\n# Load model and tokenizer\nmodel_path = '/kaggle/input/berdmodel-2'\ntokenizer_path = '/kaggle/input/berdmodel-2'\nconfig = BertConfig.from_pretrained(model_path, output_attentions=True)\nmodel = BertModel.from_pretrained(model_path, config=config)\ntokenizer = BertTokenizer.from_pretrained(tokenizer_path)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel.to(device)\n\ndef identity(x):\n    return x\n\nwith open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:\n    cpc_cv_tfidf = pickle.load(f)\n\ndef select_top_k_columns(X: Any, k: int) -> tuple[Any, np.ndarray]:\n    row_sums = X.sum(axis=0)\n    top_k_indices = np.argsort(-row_sums.A1)[:k]\n    X_top = X[:, top_k_indices]\n    return X_top, top_k_indices\n\ndef extract_important_tokens(text, model, tokenizer, num_tokens=5):\n    inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)\n    with torch.no_grad():\n        outputs = model(**inputs)\n        last_hidden_state = outputs.last_hidden_state\n        attention_weights = outputs.attentions[-1]\n        attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()\n    tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())\n    filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]\n    sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)\n    unique_important_tokens = []\n    seen_tokens = set()\n    for token in sorted_tokens:\n        if len(token[0]) > 1 and token[0] not in seen_tokens:\n            unique_important_tokens.append(token[0])\n            seen_tokens.add(token[0])\n        if len(unique_important_tokens) >= num_tokens:\n            break\n    return unique_important_tokens\n\n@dataclass\nclass Word:\n    category: str\n    content: str\n\n    def to_str(self):\n        return f\"{self.category}:{self.content}\"\n\n@dataclass\nclass State:\n    words: List[Word]\n\n    def __post_init__(self):\n        self.use = np.random.binomial(1, 0.5, len(self.words))\n\n    def to_query(self):\n        words = [word.to_str() for word, use in zip(self.words, self.use) if use]\n        if len(words) > 10000:\n            words = words[:10000]\n        return \" OR \".join(words)\n\n    def move_1(self):\n        idx = np.random.choice(len(self.words))\n        self.use[idx] = 1 - self.use[idx]\n        return self\n\nclass USPTOProblem(Annealer):\n    def __init__(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):\n        super(USPTOProblem, self).__init__(init_state)\n        self.qp = qp\n        self.searcher = searcher\n        self.target = target\n        self.Tmax = tmax\n        self.Tmin = tmin\n        self.steps = steps\n        self.max_time = max_time\n        self.copy_strategy = copy_strategy\n\n    def move(self):\n        self.state.move_1()\n\n    def energy(self):\n        query = self.state.to_query()\n        cand = whoosh_utils.execute_query(query, self.qp, self.searcher)\n        ap50_score = ap50(cand, self.target)\n        return -ap50_score\n\ndef ap50(preds: list[str], labels: list[str]) -> float:\n    precisions = []\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n        precisions.append(n_found / (e + 1))\n    return sum(precisions) / 50\n\ntest_idx = whoosh_utils.load_index(\"./test_index\")\nsearcher = whoosh_utils.get_searcher(test_idx)\nqp = whoosh_utils.get_query_parser()\n\nscores = []\nresults = []\n\nfor i in tqdm(range(len(test))):\n    target = test[i].to_numpy().flatten()[1:].tolist()\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))\n\n    if len(meta_i) == 0:\n        results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": \"ti:device\"})\n        print(\"\\t Append Dummy\", i)\n        continue\n\n    titles = meta_i.get_column(\"title\").fill_null(\"\").to_list()\n    abstracts = meta_i.get_column(\"abstract\").fill_null(\"\").to_list()\n    cpc_lists = meta_i.get_column(\"cpc\").to_list()\n    cpcs = [cpc for sublist in cpc_lists for cpc in sublist]\n\n    cpc_mat = cpc_cv_tfidf.transform(cpcs)\n    X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=3)\n    topk_cpc = cpc_cv_tfidf.get_feature_names_out()[cpc_idx]\n    topk_cpc = [Word(category=\"cpc\", content=x) for x in topk_cpc]\n\n    topk_titles = []\n    topk_abstracts = []\n    for title in titles:\n        topk_titles.extend(extract_important_tokens(title, model, tokenizer, num_tokens=10))\n    for abstract in abstracts:\n        topk_abstracts.extend(extract_important_tokens(abstract, model, tokenizer, num_tokens=30))\n\n    topk_titles = list(set(topk_titles))[:30]\n    topk_abstracts = list(set(topk_abstracts))[:30]\n    topk_words = [Word(category=\"ti\", content=x) for x in topk_titles]\n    topk_words += [Word(category=\"ab\", content=x) for x in topk_abstracts]\n    words = topk_words + topk_cpc\n    state = State(words=words)\n    problem = USPTOProblem(qp, searcher, target, state, steps=1000, max_time=10)\n    solution, score = problem.anneal()\n    print(f\"\\t Problem Number {i} Score:\", -score)\n    scores.append(-score)\n    results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": solution.to_query()})\n\nprint(\"Average Score:\", sum(scores) / len(scores))\n\n# Clean the working directory to avoid submission errors\n!rm -rf /kaggle/working/*\n\n# Generate the submission file\nsubmission = pl.DataFrame(results)\n\n# Ensure there are no missing values\nsubmission = submission.fill_null(\"\")\n\n# Save the submission file\nsubmission.write_csv(\"submission.csv\")",
      "votes": null
    },
    {
      "id": "2931355",
      "postDate": "07/21/2024 23:35:47",
      "content": "<p>Thank you for your response.<br>\nI created the attached submission file, but the same error occurred. The format also seems to be correct. What could be the problem?</p>",
      "rawMarkdown": "Thank you for your response.\nI created the attached submission file, but the same error occurred. The format also seems to be correct. What could be the problem?",
      "votes": null
    },
    {
      "id": "2931738",
      "postDate": "07/22/2024 09:43:00",
      "content": "<p>Before submitting, have some check like this. Also, run your code on a validation set and see if any of the generated queries are invalid or exceed 50 tokens.</p>\n<p>Looking at your code, you may be exceeding the limitation of 50 tokens per query (operators like OR will be also counted as tokens!). Ideally, you should change your annealing code so that it does not produce a query longer than 50 tokens.</p>\n<pre><code>queryValidator = whoosh_utils()\n\nfinal_queries = \n item  results:\n    query = item\n\n     query is None or query == :\n        final_query = \n        ()\n         ()\n\n    :\n        try:\n            queryValidator(query)\n            final_query = query\n        except:\n            final_query = \n            ()\n             ()\n\n        try:\n            token_count = whoosh_utils(query)\n            (token_count)\n             token_count &gt; :\n                final_query = \n                ()\n                 ()\n\n        except:\n            final_query = \n            ()\n\n\n    final_queries({: item, : final_query})\n</code></pre>\n<p>NB! Check also your ap50 code. Remember, if the query results exceed 50 patents, the results will be limited like this results[:50). If you are not cutting off the patents after 50, you might get a too optimistic ap50 score.</p>",
      "rawMarkdown": "Before submitting, have some check like this. Also, run your code on a validation set and see if any of the generated queries are invalid or exceed 50 tokens.\n\nLooking at your code, you may be exceeding the limitation of 50 tokens per query (operators like OR will be also counted as tokens!). Ideally, you should change your annealing code so that it does not produce a query longer than 50 tokens.\n\n```\nqueryValidator = whoosh_utils.QueryValidator()\n\nfinal_queries = []\nfor item in results:\n    query = item[\"query\"]\n    \n    if query is None or query == \"\":\n        final_query = \"ab:device\"\n        print(\"Is none\")\n        #raise Exception(\"Is none\")\n        \n    else:\n        try:\n            queryValidator.validate_query(query)\n            final_query = query\n        except:\n            final_query = \"ab:device\"\n            print(\"Not valid\")\n            #raise Exception(\"Not valid\")\n        \n        try:\n            token_count = whoosh_utils.count_query_tokens(query)\n            print(token_count)\n            if token_count > 50:\n                final_query = \"ti:mobile\"\n                print(\"Exceeds 50\")\n                #raise Exception(\"Exceeds 50\")\n\n        except:\n            final_query = \"ti:device\"\n            print(\"Error checking token count\")\n\n        \n    final_queries.append({'publication_number': item['publication_number'], 'query': final_query})\n\n```\n\nNB! Check also your ap50 code. Remember, if the query results exceed 50 patents, the results will be limited like this results[:50). If you are not cutting off the patents after 50, you might get a too optimistic ap50 score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2931198,
      "author_name": "calispohwang",
      "author_url": "",
      "post_date": "07/21/2024 18:14:17",
      "content": "<p>It seems that you are experiencing submission errors, likely due to issues in generating the submission file. Based on the provided code, here are the potential issues and the suggested fix:</p>\n<p>Potential Issues:<br>\nIncorrect Submission File Format:</p>\n<p>Your submission file may not meet the required format, which typically includes specific columns and data types.<br>\nResidual Files:</p>\n<p>There may be residual files in your working directory that are causing conflicts.<br>\nMissing Submission File Generation:</p>\n<p>It appears that the code to generate and save the submission file is missing.<br>\nSuggested Fix:<br>\nEnsure Correct Submission File Format:</p>\n<p>Verify that the submission file has the correct number of rows and columns, and that all required columns are present.<br>\nClean Working Directory:</p>\n<p>Remove any unnecessary files from the working directory to prevent conflicts.<br>\nGenerate and Save the Submission File:</p>\n<p>Add the code to create and save the submission file.<br>\nHere is the modified code with these changes:</p>\n<p>import torch<br>\nfrom transformers import BertModel, BertTokenizer, BertConfig<br>\nimport numpy as np<br>\nimport polars as pl<br>\nfrom tqdm import tqdm<br>\nimport whoosh_utils<br>\nfrom pathlib import Path<br>\nfrom dataclasses import dataclass<br>\nfrom typing import Any, List<br>\nimport random<br>\nimport signal<br>\nimport time<br>\nimport datetime<br>\nimport pickle<br>\nimport copy</p>\n<h1>Load model and tokenizer</h1>\n<p>model_path = '/kaggle/input/berdmodel-2'<br>\ntokenizer_path = '/kaggle/input/berdmodel-2'<br>\nconfig = BertConfig.from_pretrained(model_path, output_attentions=True)<br>\nmodel = BertModel.from_pretrained(model_path, config=config)<br>\ntokenizer = BertTokenizer.from_pretrained(tokenizer_path)<br>\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")<br>\nmodel.to(device)</p>\n<p>def identity(x):<br>\n    return x</p>\n<p>with open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:<br>\n    cpc_cv_tfidf = pickle.load(f)</p>\n<p>def select_top_k_columns(X: Any, k: int) -&gt; tuple[Any, np.ndarray]:<br>\n    row_sums = X.sum(axis=0)<br>\n    top_k_indices = np.argsort(-row_sums.A1)[:k]<br>\n    X_top = X[:, top_k_indices]<br>\n    return X_top, top_k_indices</p>\n<p>def extract_important_tokens(text, model, tokenizer, num_tokens=5):<br>\n    inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)<br>\n    with torch.no_grad():<br>\n        outputs = model(**inputs)<br>\n        last_hidden_state = outputs.last_hidden_state<br>\n        attention_weights = outputs.attentions[-1]<br>\n        attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()<br>\n    tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())<br>\n    filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]<br>\n    sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)<br>\n    unique_important_tokens = []<br>\n    seen_tokens = set()<br>\n    for token in sorted_tokens:<br>\n        if len(token[0]) &gt; 1 and token[0] not in seen_tokens:<br>\n            unique_important_tokens.append(token[0])<br>\n            seen_tokens.add(token[0])<br>\n        if len(unique_important_tokens) &gt;= num_tokens:<br>\n            break<br>\n    return unique_important_tokens</p>\n<p><a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\nclass Word:<br>\n    category: str<br>\n    content: str</p>\n<pre><code> ():\n     \n</code></pre>\n<p><a href=\"https://www.kaggle.com/dataclass\" target=\"_blank\">@dataclass</a><br>\nclass State:<br>\n    words: List[Word]</p>\n<pre><code> ():\n    self.use = np.random.binomial(, , (self.words))\n\n ():\n    words = [word.to_str()  word, use  (self.words, self.use)  use]\n     (words) &gt; :\n        words = words[:]\n     .join(words)\n\n ():\n    idx = np.random.choice((self.words))\n    self.use[idx] =  - self.use[idx]\n     self\n</code></pre>\n<p>class USPTOProblem(Annealer):<br>\n    def <strong>init</strong>(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):<br>\n        super(USPTOProblem, self).<strong>init</strong>(init_state)<br>\n        self.qp = qp<br>\n        self.searcher = searcher<br>\n        self.target = target<br>\n        self.Tmax = tmax<br>\n        self.Tmin = tmin<br>\n        self.steps = steps<br>\n        self.max_time = max_time<br>\n        self.copy_strategy = copy_strategy</p>\n<pre><code>def move():\n    ..move_1()\n\ndef energy():\n    query = ..to_query()\n    cand = whoosh_utils.execute_query(query, .qp, .searcher)\n    ap50_score = ap50(cand, .target)\n    return -ap50_score\n</code></pre>\n<p>def ap50(preds: list[str], labels: list[str]) -&gt; float:<br>\n    precisions = []<br>\n    n_found = 0<br>\n    for e, i in enumerate(preds):<br>\n        if i in labels:<br>\n            n_found += 1<br>\n        precisions.append(n_found / (e + 1))<br>\n    return sum(precisions) / 50</p>\n<p>test_idx = whoosh_utils.load_index(\"./test_index\")<br>\nsearcher = whoosh_utils.get_searcher(test_idx)<br>\nqp = whoosh_utils.get_query_parser()</p>\n<p>scores = []<br>\nresults = []</p>\n<p>for i in tqdm(range(len(test))):<br>\n    target = test[i].to_numpy().flatten()[1:].tolist()<br>\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))</p>\n<pre><code> (meta_i) == :\n    results({: test, : })\n    (, i)\n    continue\n\ntitles = meta_i()()()\nabstracts = meta_i()()()\ncpc_lists = meta_i()()\ncpcs = \n\ncpc_mat = cpc_cv_tfidf(cpcs)\nX_cpc, cpc_idx = (cpc_mat, k=)\ntopk_cpc = cpc_cv_tfidf()\ntopk_cpc = \n\ntopk_titles = \ntopk_abstracts = \n title  titles:\n    topk_titles((title, model, tokenizer, num_tokens=))\n abstract  abstracts:\n    topk_abstracts((abstract, model, tokenizer, num_tokens=))\n\ntopk_titles = ((topk_titles))\ntopk_abstracts = ((topk_abstracts))\ntopk_words = \ntopk_words += \nwords = topk_words + topk_cpc\nstate = (words=words)\nproblem = (qp, searcher, target, state, steps=, max_time=)\nsolution, score = problem()\n\nscores(-score)\nresults({: test, : solution()})\n</code></pre>\n<p>print(\"Average Score:\", sum(scores) / len(scores))</p>\n<h1>Clean the working directory to avoid submission errors</h1>\n<p>!rm -rf /kaggle/working/*</p>\n<h1>Generate the submission file</h1>\n<p>submission = pl.DataFrame(results)</p>\n<h1>Ensure there are no missing values</h1>\n<p>submission = submission.fill_null(\"\")</p>\n<h1>Save the submission file</h1>\n<p>submission.write_csv(\"submission.csv\")</p>",
      "votes": null,
      "replies": [
        {
          "id": 2931355,
          "author_name": "nomuraryota",
          "author_url": "",
          "post_date": "07/21/2024 23:35:47",
          "content": "<p>Thank you for your response.<br>\nI created the attached submission file, but the same error occurred. The format also seems to be correct. What could be the problem?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2931738,
              "author_name": "keakohv",
              "author_url": "",
              "post_date": "07/22/2024 09:43:00",
              "content": "<p>Before submitting, have some check like this. Also, run your code on a validation set and see if any of the generated queries are invalid or exceed 50 tokens.</p>\n<p>Looking at your code, you may be exceeding the limitation of 50 tokens per query (operators like OR will be also counted as tokens!). Ideally, you should change your annealing code so that it does not produce a query longer than 50 tokens.</p>\n<pre><code>queryValidator = whoosh_utils()\n\nfinal_queries = \n item  results:\n    query = item\n\n     query is None or query == :\n        final_query = \n        ()\n         ()\n\n    :\n        try:\n            queryValidator(query)\n            final_query = query\n        except:\n            final_query = \n            ()\n             ()\n\n        try:\n            token_count = whoosh_utils(query)\n            (token_count)\n             token_count &gt; :\n                final_query = \n                ()\n                 ()\n\n        except:\n            final_query = \n            ()\n\n\n    final_queries({: item, : final_query})\n</code></pre>\n<p>NB! Check also your ap50 code. Remember, if the query results exceed 50 patents, the results will be limited like this results[:50). If you are not cutting off the patents after 50, you might get a too optimistic ap50 score.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2930693": "I'm trying to submit using the code below, but I can't submit due to an error. Somebody please help me.\n\n\n>import torch\nfrom transformers import BertModel, BertTokenizer, BertConfig\nimport numpy as np\nimport polars as pl\nfrom tqdm import tqdm\nimport whoosh_utils\nfrom pathlib import Path\nfrom dataclasses import dataclass\nfrom typing import Any, List\nimport random\nimport signal\nimport time\nimport datetime\nimport pickle\nimport copy\nmodel_path = '/kaggle/input/berdmodel-2'\ntokenizer_path = '/kaggle/input/berdmodel-2'\nconfig = BertConfig.from_pretrained(model_path, output_attentions=True)\nmodel = BertModel.from_pretrained(model_path, config=config)\ntokenizer = BertTokenizer.from_pretrained(tokenizer_path)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel.to(device)\ndef identity(x):\n    return x\nwith open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:\n    cpc_cv_tfidf = pickle.load(f)\ndef select_top_k_columns(X: Any, k: int) -> tuple[Any, np.ndarray]:\n    # 行方向の和を計算\n    row_sums = X.sum(axis=0)\n    # 和が大きい上位k個の列のインデックスを取得\n    top_k_indices = np.argsort(-row_sums.A1)[:k]\n    # 上位k個の列を選択\n    X_top = X[:, top_k_indices]\n    return X_top, top_k_indices\ndef extract_important_tokens(text, model, tokenizer, num_tokens=5):\n    inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)\n    with torch.no_grad():\n        outputs = model(**inputs)\n        last_hidden_state = outputs.last_hidden_state\n        attention_weights = outputs.attentions[-1]  # 最後の注意層を取得\n    attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()\n    tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())\n    filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]\n    sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)\n    unique_important_tokens = []\n    seen_tokens = set()\n    for token in filtered_tokens:\n        if len(token[0]) > 1 and token[0] not in seen_tokens:\n            unique_important_tokens.append(token[0])\n            seen_tokens.add(token[0])\n        if len(unique_important_tokens) >= num_tokens:\n            break\n    return unique_important_tokens\n@dataclass\nclass Word:\n    category: str\n    content: str\n    def to_str(self):\n        return f\"{self.category}:{self.content}\"\n\n@dataclass\nclass State:\n    words: List[Word]\n\n    def __post_init__(self):\n        self.use = np.random.binomial(1, 0.5, len(self.words))\n\n    def to_query(self):\n        words = [word.to_str() for word, use in zip(self.words, self.use) if use]\n        if len(words) > 10000:  # クエリの長さを制限\n            words = words[:10000]\n        return \" OR \".join(words)\n\n    def move_1(self):\n        idx = np.random.choice(len(self.words))\n        self.use[idx] = 1 - self.use[idx]\n        return self\n\nclass USPTOProblem(Annealer):\n    def __init__(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):\n        super(USPTOProblem, self).__init__(init_state)\n        self.qp = qp\n        self.searcher = searcher\n        self.target = target\n        self.Tmax = tmax\n        self.Tmin = tmin\n        self.steps = steps\n        self.max_time = max_time\n        self.copy_strategy = copy_strategy\n\n    def move(self):\n        self.state.move_1()\n\n    def energy(self):\n        query = self.state.to_query()\n        cand = whoosh_utils.execute_query(query, self.qp, self.searcher)\n        ap50_score = ap50(cand, self.target)\n        return -ap50_score\n\ndef ap50(preds: list[str], labels: list[str]) -> float:\n    precisions = []\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n        precisions.append(n_found / (e + 1))\n    return sum(precisions) / 50\n\ntest_idx = whoosh_utils.load_index(\"./test_index\")\nsearcher = whoosh_utils.get_searcher(test_idx)\nqp = whoosh_utils.get_query_parser()\n\nscores = []\nresults = []\n\nfor i in tqdm(range(len(test))):\n    target = test[i].to_numpy().flatten()[1:].tolist()\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))\n\n    if len(meta_i) == 0:\n        results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": \"ti:device\"})\n        print(\"\\t Append Dummy\", i)\n        continue\n\n    titles = meta_i.get_column(\"title\").fill_null(\"\").to_list()\n    abstracts = meta_i.get_column(\"abstract\").fill_null(\"\").to_list()\n    cpc_lists = meta_i.get_column(\"cpc\").to_list()\n    cpcs = [cpc for sublist in cpc_lists for cpc in sublist]  # リスト内のリストを平坦化\n\n    # TF-IDF matrix for CPC codes\n    cpc_mat = cpc_cv_tfidf.transform(cpcs)\n    X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=3)\n    topk_cpc = cpc_cv_tfidf.get_feature_names_out()[cpc_idx]\n    topk_cpc = [Word(category=\"cpc\", content=x) for x in topk_cpc]\n    # Important topk words from titles and abstracts\n    topk_titles = []\n    topk_abstracts = []\n    for title in titles:\n        topk_titles.extend(extract_important_tokens(title, model, tokenizer, num_tokens=10))\n    for abstract in abstracts:\n        topk_abstracts.extend(extract_important_tokens(abstract, model, tokenizer, num_tokens=30))\n\n    # 一意にする\n    topk_titles = list(set(topk_titles))[:30]\n    topk_abstracts = list(set(topk_abstracts))[:30]\n    topk_words = [Word(category=\"ti\", content=x) for x in topk_titles]\n    topk_words += [Word(category=\"ab\", content=x) for x in topk_abstracts]\n    words = topk_words + topk_cpc\n    state = State(words=words)\n    problem = USPTOProblem(qp, searcher, target, state, steps=1000, max_time=10)\n    solution, score = problem.anneal()\n    print(f\"\\t Problem Number {i} Score:\", -score)\n    scores.append(-score)\n    results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": solution.to_query()})\nprint(\"Average Score:\", sum(scores) / len(scores))",
    "2931198": "It seems that you are experiencing submission errors, likely due to issues in generating the submission file. Based on the provided code, here are the potential issues and the suggested fix:\n\nPotential Issues:\nIncorrect Submission File Format:\n\nYour submission file may not meet the required format, which typically includes specific columns and data types.\nResidual Files:\n\nThere may be residual files in your working directory that are causing conflicts.\nMissing Submission File Generation:\n\nIt appears that the code to generate and save the submission file is missing.\nSuggested Fix:\nEnsure Correct Submission File Format:\n\nVerify that the submission file has the correct number of rows and columns, and that all required columns are present.\nClean Working Directory:\n\nRemove any unnecessary files from the working directory to prevent conflicts.\nGenerate and Save the Submission File:\n\nAdd the code to create and save the submission file.\nHere is the modified code with these changes:\n\nimport torch\nfrom transformers import BertModel, BertTokenizer, BertConfig\nimport numpy as np\nimport polars as pl\nfrom tqdm import tqdm\nimport whoosh_utils\nfrom pathlib import Path\nfrom dataclasses import dataclass\nfrom typing import Any, List\nimport random\nimport signal\nimport time\nimport datetime\nimport pickle\nimport copy\n\n# Load model and tokenizer\nmodel_path = '/kaggle/input/berdmodel-2'\ntokenizer_path = '/kaggle/input/berdmodel-2'\nconfig = BertConfig.from_pretrained(model_path, output_attentions=True)\nmodel = BertModel.from_pretrained(model_path, config=config)\ntokenizer = BertTokenizer.from_pretrained(tokenizer_path)\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel.to(device)\n\ndef identity(x):\n    return x\n\nwith open(\"/kaggle/input/uspto-ti-cpc-tfidf/cpc_cv_tfidf.pkl\", \"rb\") as f:\n    cpc_cv_tfidf = pickle.load(f)\n\ndef select_top_k_columns(X: Any, k: int) -> tuple[Any, np.ndarray]:\n    row_sums = X.sum(axis=0)\n    top_k_indices = np.argsort(-row_sums.A1)[:k]\n    X_top = X[:, top_k_indices]\n    return X_top, top_k_indices\n\ndef extract_important_tokens(text, model, tokenizer, num_tokens=5):\n    inputs = tokenizer(text, return_tensors='pt', max_length=512, truncation=True, padding='max_length').to(device)\n    with torch.no_grad():\n        outputs = model(**inputs)\n        last_hidden_state = outputs.last_hidden_state\n        attention_weights = outputs.attentions[-1]\n        attention_scores = attention_weights.mean(dim=1).squeeze(0).mean(dim=0).cpu().numpy()\n    tokens = tokenizer.convert_ids_to_tokens(inputs['input_ids'].squeeze(0).tolist())\n    filtered_tokens = [(token, score) for token, score in zip(tokens, attention_scores) if token not in tokenizer.all_special_tokens and '##' not in token]\n    sorted_tokens = sorted(filtered_tokens, key=lambda x: x[1], reverse=True)\n    unique_important_tokens = []\n    seen_tokens = set()\n    for token in sorted_tokens:\n        if len(token[0]) > 1 and token[0] not in seen_tokens:\n            unique_important_tokens.append(token[0])\n            seen_tokens.add(token[0])\n        if len(unique_important_tokens) >= num_tokens:\n            break\n    return unique_important_tokens\n\n@dataclass\nclass Word:\n    category: str\n    content: str\n\n    def to_str(self):\n        return f\"{self.category}:{self.content}\"\n\n@dataclass\nclass State:\n    words: List[Word]\n\n    def __post_init__(self):\n        self.use = np.random.binomial(1, 0.5, len(self.words))\n\n    def to_query(self):\n        words = [word.to_str() for word, use in zip(self.words, self.use) if use]\n        if len(words) > 10000:\n            words = words[:10000]\n        return \" OR \".join(words)\n\n    def move_1(self):\n        idx = np.random.choice(len(self.words))\n        self.use[idx] = 1 - self.use[idx]\n        return self\n\nclass USPTOProblem(Annealer):\n    def __init__(self, qp: Any, searcher: Any, target: List[str], init_state: State, tmax: int = 30, tmin: int = 10, steps: int = 2000, max_time: int = 10, copy_strategy: str = \"deepcopy\"):\n        super(USPTOProblem, self).__init__(init_state)\n        self.qp = qp\n        self.searcher = searcher\n        self.target = target\n        self.Tmax = tmax\n        self.Tmin = tmin\n        self.steps = steps\n        self.max_time = max_time\n        self.copy_strategy = copy_strategy\n\n    def move(self):\n        self.state.move_1()\n\n    def energy(self):\n        query = self.state.to_query()\n        cand = whoosh_utils.execute_query(query, self.qp, self.searcher)\n        ap50_score = ap50(cand, self.target)\n        return -ap50_score\n\ndef ap50(preds: list[str], labels: list[str]) -> float:\n    precisions = []\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n        precisions.append(n_found / (e + 1))\n    return sum(precisions) / 50\n\ntest_idx = whoosh_utils.load_index(\"./test_index\")\nsearcher = whoosh_utils.get_searcher(test_idx)\nqp = whoosh_utils.get_query_parser()\n\nscores = []\nresults = []\n\nfor i in tqdm(range(len(test))):\n    target = test[i].to_numpy().flatten()[1:].tolist()\n    meta_i = test_meta.filter(pl.col(\"publication_number\").is_in(target))\n\n    if len(meta_i) == 0:\n        results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": \"ti:device\"})\n        print(\"\\t Append Dummy\", i)\n        continue\n\n    titles = meta_i.get_column(\"title\").fill_null(\"\").to_list()\n    abstracts = meta_i.get_column(\"abstract\").fill_null(\"\").to_list()\n    cpc_lists = meta_i.get_column(\"cpc\").to_list()\n    cpcs = [cpc for sublist in cpc_lists for cpc in sublist]\n\n    cpc_mat = cpc_cv_tfidf.transform(cpcs)\n    X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=3)\n    topk_cpc = cpc_cv_tfidf.get_feature_names_out()[cpc_idx]\n    topk_cpc = [Word(category=\"cpc\", content=x) for x in topk_cpc]\n\n    topk_titles = []\n    topk_abstracts = []\n    for title in titles:\n        topk_titles.extend(extract_important_tokens(title, model, tokenizer, num_tokens=10))\n    for abstract in abstracts:\n        topk_abstracts.extend(extract_important_tokens(abstract, model, tokenizer, num_tokens=30))\n\n    topk_titles = list(set(topk_titles))[:30]\n    topk_abstracts = list(set(topk_abstracts))[:30]\n    topk_words = [Word(category=\"ti\", content=x) for x in topk_titles]\n    topk_words += [Word(category=\"ab\", content=x) for x in topk_abstracts]\n    words = topk_words + topk_cpc\n    state = State(words=words)\n    problem = USPTOProblem(qp, searcher, target, state, steps=1000, max_time=10)\n    solution, score = problem.anneal()\n    print(f\"\\t Problem Number {i} Score:\", -score)\n    scores.append(-score)\n    results.append({\"publication_number\": test[i, \"publication_number\"], \"query\": solution.to_query()})\n\nprint(\"Average Score:\", sum(scores) / len(scores))\n\n# Clean the working directory to avoid submission errors\n!rm -rf /kaggle/working/*\n\n# Generate the submission file\nsubmission = pl.DataFrame(results)\n\n# Ensure there are no missing values\nsubmission = submission.fill_null(\"\")\n\n# Save the submission file\nsubmission.write_csv(\"submission.csv\")",
    "2931355": "Thank you for your response.\nI created the attached submission file, but the same error occurred. The format also seems to be correct. What could be the problem?",
    "2931738": "Before submitting, have some check like this. Also, run your code on a validation set and see if any of the generated queries are invalid or exceed 50 tokens.\n\nLooking at your code, you may be exceeding the limitation of 50 tokens per query (operators like OR will be also counted as tokens!). Ideally, you should change your annealing code so that it does not produce a query longer than 50 tokens.\n\n```\nqueryValidator = whoosh_utils.QueryValidator()\n\nfinal_queries = []\nfor item in results:\n    query = item[\"query\"]\n    \n    if query is None or query == \"\":\n        final_query = \"ab:device\"\n        print(\"Is none\")\n        #raise Exception(\"Is none\")\n        \n    else:\n        try:\n            queryValidator.validate_query(query)\n            final_query = query\n        except:\n            final_query = \"ab:device\"\n            print(\"Not valid\")\n            #raise Exception(\"Not valid\")\n        \n        try:\n            token_count = whoosh_utils.count_query_tokens(query)\n            print(token_count)\n            if token_count > 50:\n                final_query = \"ti:mobile\"\n                print(\"Exceeds 50\")\n                #raise Exception(\"Exceeds 50\")\n\n        except:\n            final_query = \"ti:device\"\n            print(\"Error checking token count\")\n\n        \n    final_queries.append({'publication_number': item['publication_number'], 'query': final_query})\n\n```\n\nNB! Check also your ap50 code. Remember, if the query results exceed 50 patents, the results will be limited like this results[:50). If you are not cutting off the patents after 50, you might get a too optimistic ap50 score."
  },
  "source": "meta"
}