{
  "id": 522380,
  "title": "18th place solution",
  "url": "/competitions/uspto-explainable-ai/discussion/522380",
  "author_name": "Raki",
  "post_date": "2024-07-25T21:43:30.076000",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>Overview of the Approach</h1>\n<p>Our final approach employed a preprocessing and filtering pipeline designed to minimize overfitting. Initially, we utilized a complex method involving CPC code searches with linear programming, multiple algorithms including decision trees, and logical chains (NOT, OR). Ultimately, we simplified this to a straightforward OR-chain based on single word matches. We created an <a href=\"https://www.kaggle.com/datasets/raki21/uspto-full-error-corrected\" target=\"_blank\">SQLite database</a> containing filtered information for all publications, implemented in this <a href=\"https://www.kaggle.com/code/raki21/simple-or-chain-sqlite-magic-free\" target=\"_blank\">notebook using only a single OR-chain</a> to achieve a 0.72+ score (our best submission was only 0.66, this notebook just changed the target score of our approach from precision to mAP, we explain the confusion below).</p>\n<h3>Details</h3>\n<ol>\n<li><p><strong>SQL Database Creation</strong>:</p>\n<ul>\n<li>Data was split by words to create an SQL database for inference.</li>\n<li>Each column was processed using <code>whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()</code>, followed by removing duplicates to save memory and because single word search was initially prioritized.</li></ul></li>\n<li><p><strong>Neighbors at Inference</strong>:</p>\n<ul>\n<li>We loaded the 125k main publications and their neighbors into a dictionary for O(1) retrieval during inference. Additional neighbors-neighbors contributed to a word occurrence dictionary but were not used directly.</li></ul></li>\n<li><p><strong>Generation of Feature DataFrame</strong>:</p>\n<ul>\n<li>We identified all words occurring in any neighbor and added them to the feature list. A dataframe was created for each patent with all features as columns and 50 rows indicating occurrence in neighbors, resulting in approximately 10,000 unique words.</li></ul></li>\n<li><p><strong>OR-Chain Generation</strong>:</p>\n<ul>\n<li>Features were iteratively selected based on a function penalizing false positives and rewarding true matches. Parameters for penalizing false positives were optimized to select the best scoring OR-chain, focusing on precision rather than mean Average Precision (mAP).</li></ul></li>\n</ol>\n<h3>Problem: Metric Misinterpretation</h3>\n<p>A significant oversight was the misinterpretation of the scoring metric. Initially, we assumed the metric could be treated as precision since false matches filled up to 50 without considering order. Later, we realized that tf-idf and ordering were used but did not reevaluate our precision-based approach. This led to opting for a 35/15 split instead of a 25/0 split, resulting in lower scores (0.7 vs 0.84).</p>\n<h3>Missing Additions</h3>\n<p>Potential improvements included:</p>\n<ul>\n<li>Faster inference with Ray/GPU to utilize more neighbors.</li>\n<li>Usage of AND for multiple OR-chains.</li>\n<li>A semantically independent NOT-OR chain, leveraging a global word count dictionary to omit highly frequent words missing from all positive cases. This would also need some way to identify omissions that have lots of internal overlap, maybe some semantic similarity measure, as often you would have top 3 omissions being something like [material, materials, metal].</li>\n</ul>\n<h3>Final Remarks</h3>\n<p>We extend our gratitude to the competition organizers, participants who shared their approaches and ideas, and congratulations to the winners!</p>",
  "messages": [
    {
      "id": 2936167,
      "postDate": "2024-07-25T21:43:30.077Z",
      "content": "<h1>Overview of the Approach</h1>\n<p>Our final approach employed a preprocessing and filtering pipeline designed to minimize overfitting. Initially, we utilized a complex method involving CPC code searches with linear programming, multiple algorithms including decision trees, and logical chains (NOT, OR). Ultimately, we simplified this to a straightforward OR-chain based on single word matches. We created an <a href=\"https://www.kaggle.com/datasets/raki21/uspto-full-error-corrected\" target=\"_blank\">SQLite database</a> containing filtered information for all publications, implemented in this <a href=\"https://www.kaggle.com/code/raki21/simple-or-chain-sqlite-magic-free\" target=\"_blank\">notebook using only a single OR-chain</a> to achieve a 0.72+ score (our best submission was only 0.66, this notebook just changed the target score of our approach from precision to mAP, we explain the confusion below).</p>\n<h3>Details</h3>\n<ol>\n<li><p><strong>SQL Database Creation</strong>:</p>\n<ul>\n<li>Data was split by words to create an SQL database for inference.</li>\n<li>Each column was processed using <code>whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()</code>, followed by removing duplicates to save memory and because single word search was initially prioritized.</li></ul></li>\n<li><p><strong>Neighbors at Inference</strong>:</p>\n<ul>\n<li>We loaded the 125k main publications and their neighbors into a dictionary for O(1) retrieval during inference. Additional neighbors-neighbors contributed to a word occurrence dictionary but were not used directly.</li></ul></li>\n<li><p><strong>Generation of Feature DataFrame</strong>:</p>\n<ul>\n<li>We identified all words occurring in any neighbor and added them to the feature list. A dataframe was created for each patent with all features as columns and 50 rows indicating occurrence in neighbors, resulting in approximately 10,000 unique words.</li></ul></li>\n<li><p><strong>OR-Chain Generation</strong>:</p>\n<ul>\n<li>Features were iteratively selected based on a function penalizing false positives and rewarding true matches. Parameters for penalizing false positives were optimized to select the best scoring OR-chain, focusing on precision rather than mean Average Precision (mAP).</li></ul></li>\n</ol>\n<h3>Problem: Metric Misinterpretation</h3>\n<p>A significant oversight was the misinterpretation of the scoring metric. Initially, we assumed the metric could be treated as precision since false matches filled up to 50 without considering order. Later, we realized that tf-idf and ordering were used but did not reevaluate our precision-based approach. This led to opting for a 35/15 split instead of a 25/0 split, resulting in lower scores (0.7 vs 0.84).</p>\n<h3>Missing Additions</h3>\n<p>Potential improvements included:</p>\n<ul>\n<li>Faster inference with Ray/GPU to utilize more neighbors.</li>\n<li>Usage of AND for multiple OR-chains.</li>\n<li>A semantically independent NOT-OR chain, leveraging a global word count dictionary to omit highly frequent words missing from all positive cases. This would also need some way to identify omissions that have lots of internal overlap, maybe some semantic similarity measure, as often you would have top 3 omissions being something like [material, materials, metal].</li>\n</ul>\n<h3>Final Remarks</h3>\n<p>We extend our gratitude to the competition organizers, participants who shared their approaches and ideas, and congratulations to the winners!</p>",
      "rawMarkdown": "# Overview of the Approach\n\nOur final approach employed a preprocessing and filtering pipeline designed to minimize overfitting. Initially, we utilized a complex method involving CPC code searches with linear programming, multiple algorithms including decision trees, and logical chains (NOT, OR). Ultimately, we simplified this to a straightforward OR-chain based on single word matches. We created an [SQLite database](https://www.kaggle.com/datasets/raki21/uspto-full-error-corrected) containing filtered information for all publications, implemented in this [notebook using only a single OR-chain](https://www.kaggle.com/code/raki21/simple-or-chain-sqlite-magic-free) to achieve a 0.72+ score (our best submission was only 0.66, this notebook just changed the target score of our approach from precision to mAP, we explain the confusion below).\n\n### Details\n\n0. **SQL Database Creation**:\n    - Data was split by words to create an SQL database for inference.\n    - Each column was processed using `whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()`, followed by removing duplicates to save memory and because single word search was initially prioritized.\n\n1. **Neighbors at Inference**:\n    - We loaded the 125k main publications and their neighbors into a dictionary for O(1) retrieval during inference. Additional neighbors-neighbors contributed to a word occurrence dictionary but were not used directly.\n\n2. **Generation of Feature DataFrame**:\n    - We identified all words occurring in any neighbor and added them to the feature list. A dataframe was created for each patent with all features as columns and 50 rows indicating occurrence in neighbors, resulting in approximately 10,000 unique words.\n\n3. **OR-Chain Generation**:\n    - Features were iteratively selected based on a function penalizing false positives and rewarding true matches. Parameters for penalizing false positives were optimized to select the best scoring OR-chain, focusing on precision rather than mean Average Precision (mAP).\n\n### Problem: Metric Misinterpretation\nA significant oversight was the misinterpretation of the scoring metric. Initially, we assumed the metric could be treated as precision since false matches filled up to 50 without considering order. Later, we realized that tf-idf and ordering were used but did not reevaluate our precision-based approach. This led to opting for a 35/15 split instead of a 25/0 split, resulting in lower scores (0.7 vs 0.84).\n\n### Missing Additions\nPotential improvements included:\n- Faster inference with Ray/GPU to utilize more neighbors.\n- Usage of AND for multiple OR-chains.\n- A semantically independent NOT-OR chain, leveraging a global word count dictionary to omit highly frequent words missing from all positive cases. This would also need some way to identify omissions that have lots of internal overlap, maybe some semantic similarity measure, as often you would have top 3 omissions being something like [material, materials, metal].\n\n### Final Remarks\nWe extend our gratitude to the competition organizers, participants who shared their approaches and ideas, and congratulations to the winners!\n",
      "votes": 9
    },
    {
      "id": 2936361,
      "postDate": "2024-07-26T05:01:07.700Z",
      "content": "<p>Thank you for sharing your detailed approach and insights; it's fascinating to see how you tackled the challenge and addressed metric misinterpretation—congratulations on your innovative solution! <a href=\"https://www.kaggle.com/raki21\" target=\"_blank\">@raki21</a> </p>",
      "rawMarkdown": "Thank you for sharing your detailed approach and insights; it's fascinating to see how you tackled the challenge and addressed metric misinterpretation—congratulations on your innovative solution! @raki21 \n\n",
      "votes": 1
    },
    {
      "id": 2936286,
      "postDate": "2024-07-26T03:32:52.440Z",
      "content": "<p>Congrats on your achievement as a non-magician, <a href=\"https://www.kaggle.com/raki21\" target=\"_blank\">@raki21</a>, and thanks for sharing all the details. Can you elaborate on what you mean by \"opting for a 35/15 split instead of a 25/0 split\"?</p>",
      "rawMarkdown": "Congrats on your achievement as a non-magician, @raki21, and thanks for sharing all the details. Can you elaborate on what you mean by \"opting for a 35/15 split instead of a 25/0 split\"?",
      "votes": 1,
      "replies": [
        {
          "id": 2936560,
          "postDate": "2024-07-26T09:08:40.983Z",
          "content": "<p>Basically we optimized in a way that would choose a query yielding 35 true positives (part of the 50 real neighbors) and 15 false positives (mismatched patents) as better then the 25 true positive, 0 false positive match case, as the 50 matches are filled up anyway leading to 25 true positives and 25 false positives. Because there is an ordering though and the fill-up of mismatches happens to the end of the prediction list, the 25/0 case is better. If ordering was random mAP ~ precision and 35/15 vs 25/0 -&gt; 25/25 gets 0.7 vs 0.5 score. With actual ordering we get 0.7 vs 0.84 though, making the query avoiding false positives much stronger. </p>",
          "rawMarkdown": "Basically we optimized in a way that would choose a query yielding 35 true positives (part of the 50 real neighbors) and 15 false positives (mismatched patents) as better then the 25 true positive, 0 false positive match case, as the 50 matches are filled up anyway leading to 25 true positives and 25 false positives. Because there is an ordering though and the fill-up of mismatches happens to the end of the prediction list, the 25/0 case is better. If ordering was random mAP ~ precision and 35/15 vs 25/0 -> 25/25 gets 0.7 vs 0.5 score. With actual ordering we get 0.7 vs 0.84 though, making the query avoiding false positives much stronger. ",
          "votes": 1,
          "replies": [
            {
              "id": 2936657,
              "postDate": "2024-07-26T11:16:40.897Z",
              "content": "<p><a href=\"https://www.kaggle.com/raki21\" target=\"_blank\">@raki21</a> I tried the 30/0 approach, but it didn't work for me, possibly due to some missing logic. Thanks for your explanation!</p>",
              "rawMarkdown": "@raki21 I tried the 30/0 approach, but it didn't work for me, possibly due to some missing logic. Thanks for your explanation!"
            }
          ]
        }
      ]
    },
    {
      "id": 2936285,
      "postDate": "2024-07-26T03:31:14.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2946950,
      "postDate": "2024-08-04T19:54:43.823Z",
      "content": "<p>Thanks for suggesting to do this :)</p>",
      "rawMarkdown": "Thanks for suggesting to do this :)"
    }
  ],
  "comments": [
    {
      "id": 2936361,
      "author_name": "Aadit Shukla",
      "author_url": "",
      "post_date": "2024-07-26T05:01:07.700000",
      "content": "<p>Thank you for sharing your detailed approach and insights; it's fascinating to see how you tackled the challenge and addressed metric misinterpretation—congratulations on your innovative solution! <a href=\"https://www.kaggle.com/raki21\" target=\"_blank\">@raki21</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2936286,
      "author_name": "FullEmpty",
      "author_url": "",
      "post_date": "2024-07-26T03:32:52.440000",
      "content": "<p>Congrats on your achievement as a non-magician, <a href=\"https://www.kaggle.com/raki21\" target=\"_blank\">@raki21</a>, and thanks for sharing all the details. Can you elaborate on what you mean by \"opting for a 35/15 split instead of a 25/0 split\"?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2936560,
          "author_name": "Raki",
          "author_url": "",
          "post_date": "2024-07-26T09:08:40.983000",
          "content": "<p>Basically we optimized in a way that would choose a query yielding 35 true positives (part of the 50 real neighbors) and 15 false positives (mismatched patents) as better then the 25 true positive, 0 false positive match case, as the 50 matches are filled up anyway leading to 25 true positives and 25 false positives. Because there is an ordering though and the fill-up of mismatches happens to the end of the prediction list, the 25/0 case is better. If ordering was random mAP ~ precision and 35/15 vs 25/0 -&gt; 25/25 gets 0.7 vs 0.5 score. With actual ordering we get 0.7 vs 0.84 though, making the query avoiding false positives much stronger. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2936657,
              "author_name": "FullEmpty",
              "author_url": "",
              "post_date": "2024-07-26T11:16:40.897000",
              "content": "<p><a href=\"https://www.kaggle.com/raki21\" target=\"_blank\">@raki21</a> I tried the 30/0 approach, but it didn't work for me, possibly due to some missing logic. Thanks for your explanation!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2936285,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-07-26T03:31:14.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2946950,
      "author_name": "Lukas Bogenrieder",
      "author_url": "",
      "post_date": "2024-08-04T19:54:43.823000",
      "content": "<p>Thanks for suggesting to do this :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2936167": "# Overview of the Approach\n\nOur final approach employed a preprocessing and filtering pipeline designed to minimize overfitting. Initially, we utilized a complex method involving CPC code searches with linear programming, multiple algorithms including decision trees, and logical chains (NOT, OR). Ultimately, we simplified this to a straightforward OR-chain based on single word matches. We created an [SQLite database](https://www.kaggle.com/datasets/raki21/uspto-full-error-corrected) containing filtered information for all publications, implemented in this [notebook using only a single OR-chain](https://www.kaggle.com/code/raki21/simple-or-chain-sqlite-magic-free) to achieve a 0.72+ score (our best submission was only 0.66, this notebook just changed the target score of our approach from precision to mAP, we explain the confusion below).\n\n### Details\n\n0. **SQL Database Creation**:\n    - Data was split by words to create an SQL database for inference.\n    - Each column was processed using `whoosh.analysis.StandardAnalyzer(stoplist=BRS_STOPWORDS) | NumberFilter()`, followed by removing duplicates to save memory and because single word search was initially prioritized.\n\n1. **Neighbors at Inference**:\n    - We loaded the 125k main publications and their neighbors into a dictionary for O(1) retrieval during inference. Additional neighbors-neighbors contributed to a word occurrence dictionary but were not used directly.\n\n2. **Generation of Feature DataFrame**:\n    - We identified all words occurring in any neighbor and added them to the feature list. A dataframe was created for each patent with all features as columns and 50 rows indicating occurrence in neighbors, resulting in approximately 10,000 unique words.\n\n3. **OR-Chain Generation**:\n    - Features were iteratively selected based on a function penalizing false positives and rewarding true matches. Parameters for penalizing false positives were optimized to select the best scoring OR-chain, focusing on precision rather than mean Average Precision (mAP).\n\n### Problem: Metric Misinterpretation\nA significant oversight was the misinterpretation of the scoring metric. Initially, we assumed the metric could be treated as precision since false matches filled up to 50 without considering order. Later, we realized that tf-idf and ordering were used but did not reevaluate our precision-based approach. This led to opting for a 35/15 split instead of a 25/0 split, resulting in lower scores (0.7 vs 0.84).\n\n### Missing Additions\nPotential improvements included:\n- Faster inference with Ray/GPU to utilize more neighbors.\n- Usage of AND for multiple OR-chains.\n- A semantically independent NOT-OR chain, leveraging a global word count dictionary to omit highly frequent words missing from all positive cases. This would also need some way to identify omissions that have lots of internal overlap, maybe some semantic similarity measure, as often you would have top 3 omissions being something like [material, materials, metal].\n\n### Final Remarks\nWe extend our gratitude to the competition organizers, participants who shared their approaches and ideas, and congratulations to the winners!\n",
    "2936361": "Thank you for sharing your detailed approach and insights; it's fascinating to see how you tackled the challenge and addressed metric misinterpretation—congratulations on your innovative solution! @raki21 \n\n",
    "2936286": "Congrats on your achievement as a non-magician, @raki21, and thanks for sharing all the details. Can you elaborate on what you mean by \"opting for a 35/15 split instead of a 25/0 split\"?",
    "2936285": "",
    "2946950": "Thanks for suggesting to do this :)"
  }
}