{
  "id": 522205,
  "title": "47th place solution",
  "url": "/competitions/uspto-explainable-ai/discussion/522205",
  "author_name": "HayatoFujihara",
  "post_date": "2024-07-25T00:35:20.366000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thank you to the hosts and participants for hosting this competition. I learned a lot.</p>\n<p>My solution is as follows.</p>\n<h3>Query check for single word and CPC</h3>\n<p>I checked the relevance by running queries for the words and CPC obtained by tfidf. This increased the score by about 0.04.</p>\n<pre><code>     j, word  (topk_words):\n        ti_query =  + word\n        cand = whoosh_utils.execute_query(ti_query, qp, searcher)\n        ti_score = ap50_true(cand, target)\n        ti_score += ((topk_words) - j) * \n        ti_scores.append(ti_score)\n\n     j, cpc  (topk_cpc):\n        cpc_query =  + cpc\n        cand = whoosh_utils.execute_query(cpc_query, qp, searcher)\n        cpc_score = ap50_true(cand, target)\n        cpc_score += ((topk_cpc) - j) * \n        cpc_scores.append(cpc_score)\n</code></pre>\n<h3>Difficulty assessment</h3>\n<p>When meta_i was divided into 5 parts and the CPCs obtained by tfidf were all the same, the score tended to drop significantly.</p>\n<p>Perhaps there was an adjacent patent that did not have a CPC.</p>\n<p>When that condition was met, a two-word search was added to give more importance to the word.</p>\n<pre><code>    meta_i_list = []\n     j  ():\n        start_index = j*\n        end_index = (start_index + , (meta_i))\n         start_index &gt;= (meta_i):\n            \n    meta_i_list.append(meta_i[start_index:end_index])\n\n    cpc_mat_list_d = [cpc_cv_tfidf.transform(m.get_column())  m  meta_i_list]\n    cpc_idx_list_d = []\n     cpc_mat_d  cpc_mat_list_d:\n        X_cpc_d, cpc_idx_d = select_top_k_columns(cpc_mat_d, k=)\n        cpc_idx_list_d.append(cpc_idx_d)\n    cpc_idx_list_d = np.unique(cpc_idx_list_d)\n    ((cpc_idx_list_d))\n    difficulty = \n     (cpc_idx_list_d) &lt;= :\n        difficulty = \n\n     difficulty:\n        X_ti, idx = select_top_k_columns(ti_mat, k=)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=)\n    :\n        X_ti, idx = select_top_k_columns(ti_mat, k=)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=)\n</code></pre>\n<h3>Random Judgment</h3>\n<p>The area that had to be explored was so large that random judgment would have helped improve the score.</p>\n<pre><code>     ():\n        p =  +  * np.random.choice(())\n        self.use = np.random.binomial(, p, (self.words))\n         (self.words) &gt;=   np.count_nonzero(self.use == ) == :        \n            self.use = np.random.binomial(, p, (self.words))\n\n         self\n</code></pre>",
  "messages": [
    {
      "id": 2935072,
      "postDate": "2024-07-25T00:35:20.367Z",
      "content": "<p>Thank you to the hosts and participants for hosting this competition. I learned a lot.</p>\n<p>My solution is as follows.</p>\n<h3>Query check for single word and CPC</h3>\n<p>I checked the relevance by running queries for the words and CPC obtained by tfidf. This increased the score by about 0.04.</p>\n<pre><code>     j, word  (topk_words):\n        ti_query =  + word\n        cand = whoosh_utils.execute_query(ti_query, qp, searcher)\n        ti_score = ap50_true(cand, target)\n        ti_score += ((topk_words) - j) * \n        ti_scores.append(ti_score)\n\n     j, cpc  (topk_cpc):\n        cpc_query =  + cpc\n        cand = whoosh_utils.execute_query(cpc_query, qp, searcher)\n        cpc_score = ap50_true(cand, target)\n        cpc_score += ((topk_cpc) - j) * \n        cpc_scores.append(cpc_score)\n</code></pre>\n<h3>Difficulty assessment</h3>\n<p>When meta_i was divided into 5 parts and the CPCs obtained by tfidf were all the same, the score tended to drop significantly.</p>\n<p>Perhaps there was an adjacent patent that did not have a CPC.</p>\n<p>When that condition was met, a two-word search was added to give more importance to the word.</p>\n<pre><code>    meta_i_list = []\n     j  ():\n        start_index = j*\n        end_index = (start_index + , (meta_i))\n         start_index &gt;= (meta_i):\n            \n    meta_i_list.append(meta_i[start_index:end_index])\n\n    cpc_mat_list_d = [cpc_cv_tfidf.transform(m.get_column())  m  meta_i_list]\n    cpc_idx_list_d = []\n     cpc_mat_d  cpc_mat_list_d:\n        X_cpc_d, cpc_idx_d = select_top_k_columns(cpc_mat_d, k=)\n        cpc_idx_list_d.append(cpc_idx_d)\n    cpc_idx_list_d = np.unique(cpc_idx_list_d)\n    ((cpc_idx_list_d))\n    difficulty = \n     (cpc_idx_list_d) &lt;= :\n        difficulty = \n\n     difficulty:\n        X_ti, idx = select_top_k_columns(ti_mat, k=)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=)\n    :\n        X_ti, idx = select_top_k_columns(ti_mat, k=)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=)\n</code></pre>\n<h3>Random Judgment</h3>\n<p>The area that had to be explored was so large that random judgment would have helped improve the score.</p>\n<pre><code>     ():\n        p =  +  * np.random.choice(())\n        self.use = np.random.binomial(, p, (self.words))\n         (self.words) &gt;=   np.count_nonzero(self.use == ) == :        \n            self.use = np.random.binomial(, p, (self.words))\n\n         self\n</code></pre>",
      "rawMarkdown": "Thank you to the hosts and participants for hosting this competition. I learned a lot.\n\nMy solution is as follows.\n\n### Query check for single word and CPC\n\nI checked the relevance by running queries for the words and CPC obtained by tfidf. This increased the score by about 0.04.\n\n```python\n    for j, word in enumerate(topk_words):\n        ti_query = f\"ti:\" + word\n        cand = whoosh_utils.execute_query(ti_query, qp, searcher)\n        ti_score = ap50_true(cand, target)\n        ti_score += (len(topk_words) - j) * 0.00001\n        ti_scores.append(ti_score)\n\n    for j, cpc in enumerate(topk_cpc):\n        cpc_query = f\"cpc:\" + cpc\n        cand = whoosh_utils.execute_query(cpc_query, qp, searcher)\n        cpc_score = ap50_true(cand, target)\n        cpc_score += (len(topk_cpc) - j) * 0.00001\n        cpc_scores.append(cpc_score)\n```\n\n### Difficulty assessment\n\nWhen meta_i was divided into 5 parts and the CPCs obtained by tfidf were all the same, the score tended to drop significantly.\n\nPerhaps there was an adjacent patent that did not have a CPC.\n\nWhen that condition was met, a two-word search was added to give more importance to the word.\n\n```python\n    meta_i_list = []\n    for j in range(5):\n        start_index = j*10\n        end_index = min(start_index + 10, len(meta_i))\n        if start_index >= len(meta_i):\n            break\n    meta_i_list.append(meta_i[start_index:end_index])\n\n    cpc_mat_list_d = [cpc_cv_tfidf.transform(m.get_column(\"cpc\")) for m in meta_i_list]\n    cpc_idx_list_d = []\n    for cpc_mat_d in cpc_mat_list_d:\n        X_cpc_d, cpc_idx_d = select_top_k_columns(cpc_mat_d, k=4)\n        cpc_idx_list_d.append(cpc_idx_d)\n    cpc_idx_list_d = np.unique(cpc_idx_list_d)\n    print(len(cpc_idx_list_d))\n    difficulty = False\n    if len(cpc_idx_list_d) <= 4:\n        difficulty = True\n\n    if difficulty:\n        X_ti, idx = select_top_k_columns(ti_mat, k=100)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=30)\n    else:\n        X_ti, idx = select_top_k_columns(ti_mat, k=30)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=50)\n```\n\n### Random Judgment\n\nThe area that had to be explored was so large that random judgment would have helped improve the score.\n\n```python\n    def move_random(self):\n        p = 0.65 + 0.05 * np.random.choice(range(6))\n        self.use = np.random.binomial(1, p, len(self.words))\n        while len(self.words) >= 1 and np.count_nonzero(self.use == 1) == 0:        \n            self.use = np.random.binomial(1, p, len(self.words))\n        \n        return self\n```",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2935072": "Thank you to the hosts and participants for hosting this competition. I learned a lot.\n\nMy solution is as follows.\n\n### Query check for single word and CPC\n\nI checked the relevance by running queries for the words and CPC obtained by tfidf. This increased the score by about 0.04.\n\n```python\n    for j, word in enumerate(topk_words):\n        ti_query = f\"ti:\" + word\n        cand = whoosh_utils.execute_query(ti_query, qp, searcher)\n        ti_score = ap50_true(cand, target)\n        ti_score += (len(topk_words) - j) * 0.00001\n        ti_scores.append(ti_score)\n\n    for j, cpc in enumerate(topk_cpc):\n        cpc_query = f\"cpc:\" + cpc\n        cand = whoosh_utils.execute_query(cpc_query, qp, searcher)\n        cpc_score = ap50_true(cand, target)\n        cpc_score += (len(topk_cpc) - j) * 0.00001\n        cpc_scores.append(cpc_score)\n```\n\n### Difficulty assessment\n\nWhen meta_i was divided into 5 parts and the CPCs obtained by tfidf were all the same, the score tended to drop significantly.\n\nPerhaps there was an adjacent patent that did not have a CPC.\n\nWhen that condition was met, a two-word search was added to give more importance to the word.\n\n```python\n    meta_i_list = []\n    for j in range(5):\n        start_index = j*10\n        end_index = min(start_index + 10, len(meta_i))\n        if start_index >= len(meta_i):\n            break\n    meta_i_list.append(meta_i[start_index:end_index])\n\n    cpc_mat_list_d = [cpc_cv_tfidf.transform(m.get_column(\"cpc\")) for m in meta_i_list]\n    cpc_idx_list_d = []\n    for cpc_mat_d in cpc_mat_list_d:\n        X_cpc_d, cpc_idx_d = select_top_k_columns(cpc_mat_d, k=4)\n        cpc_idx_list_d.append(cpc_idx_d)\n    cpc_idx_list_d = np.unique(cpc_idx_list_d)\n    print(len(cpc_idx_list_d))\n    difficulty = False\n    if len(cpc_idx_list_d) <= 4:\n        difficulty = True\n\n    if difficulty:\n        X_ti, idx = select_top_k_columns(ti_mat, k=100)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=30)\n    else:\n        X_ti, idx = select_top_k_columns(ti_mat, k=30)\n        X_cpc, cpc_idx = select_top_k_columns(cpc_mat, k=50)\n```\n\n### Random Judgment\n\nThe area that had to be explored was so large that random judgment would have helped improve the score.\n\n```python\n    def move_random(self):\n        p = 0.65 + 0.05 * np.random.choice(range(6))\n        self.use = np.random.binomial(1, p, len(self.words))\n        while len(self.words) >= 1 and np.count_nonzero(self.use == 1) == 0:        \n            self.use = np.random.binomial(1, p, len(self.words))\n        \n        return self\n```"
  }
}