{
  "id": 522327,
  "title": "22nd Place Solution - Simple TF-IDF",
  "url": "/competitions/uspto-explainable-ai/writeups/hajime-tamura-22nd-place-solution-simple-tf-idf",
  "author_name": "",
  "post_date": "2024-07-26T05:31:55.037Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thank you for organizing this very interesting competition. And thank you to all participants.<br>\nThe 22nd place solution is a simple solution using only “title” and “cpc”.</p>\n<h1>Creating candidates</h1>\n<h3>Step 1: 90 candidates by TF-IDF (title)</h3>\n<ul>\n<li>stopwords='english'</li>\n<li>ngram=(1, 3)</li>\n<li>max_df=100</li>\n<li>min_df=3</li>\n<li>If a word is in another word set, exclude it</li>\n<li>Exclude words with less than 5 letters</li>\n</ul>\n<p>Step 1 may not reach 60 because the title is too short. In that case, add Step 2 to the candidates.</p>\n<h3>Step 2: 60 candidates by TF-IDF (cpc)</h3>\n<ul>\n<li>max_df=0.01</li>\n</ul>\n<p>The sum of candidates in Step 1 and Step 2 may not reach 30 because some publication_numbers have very little information on title and cpc. In that case, add Step 3 to the candidates.</p>\n<h3>Step 3: 10 candidates by word occurrence of title</h3>\n<ul>\n<li>Split title into words and sorted by number of occurrences (max 10)</li>\n</ul>\n<h1>Annealing algorithm</h1>\n<p>Using these candidates, run the Annealing algorithm. However, I am unable to understand the annealing algorithm and use it directly from the following two excellent public notebooks.<br>\n<a href=\"https://www.kaggle.com/tubotubo\" target=\"_blank\">@tubotubo</a> <a href=\"https://www.kaggle.com/andrey67\" target=\"_blank\">@andrey67</a> I really appreciate it.</p>\n<p><a href=\"https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline\" target=\"_blank\">https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline</a><br>\n<a href=\"https://www.kaggle.com/code/andrey67/uspto-annealing-lb-0-31\" target=\"_blank\">https://www.kaggle.com/code/andrey67/uspto-annealing-lb-0-31</a></p>\n<p>Of course, I could not find MAGIC at all. It is very amazing who found it!</p>",
  "messages": [
    {
      "id": "2935831",
      "postDate": "07/25/2024 15:18:55",
      "content": "<p>Thank you for organizing this very interesting competition. And thank you to all participants.<br>\nThe 22nd place solution is a simple solution using only “title” and “cpc”.</p>\n<h1>Creating candidates</h1>\n<h3>Step 1: 90 candidates by TF-IDF (title)</h3>\n<ul>\n<li>stopwords='english'</li>\n<li>ngram=(1, 3)</li>\n<li>max_df=100</li>\n<li>min_df=3</li>\n<li>If a word is in another word set, exclude it</li>\n<li>Exclude words with less than 5 letters</li>\n</ul>\n<p>Step 1 may not reach 60 because the title is too short. In that case, add Step 2 to the candidates.</p>\n<h3>Step 2: 60 candidates by TF-IDF (cpc)</h3>\n<ul>\n<li>max_df=0.01</li>\n</ul>\n<p>The sum of candidates in Step 1 and Step 2 may not reach 30 because some publication_numbers have very little information on title and cpc. In that case, add Step 3 to the candidates.</p>\n<h3>Step 3: 10 candidates by word occurrence of title</h3>\n<ul>\n<li>Split title into words and sorted by number of occurrences (max 10)</li>\n</ul>\n<h1>Annealing algorithm</h1>\n<p>Using these candidates, run the Annealing algorithm. However, I am unable to understand the annealing algorithm and use it directly from the following two excellent public notebooks.<br>\n<a href=\"https://www.kaggle.com/tubotubo\" target=\"_blank\">@tubotubo</a> <a href=\"https://www.kaggle.com/andrey67\" target=\"_blank\">@andrey67</a> I really appreciate it.</p>\n<p><a href=\"https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline\" target=\"_blank\">https://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline</a><br>\n<a href=\"https://www.kaggle.com/code/andrey67/uspto-annealing-lb-0-31\" target=\"_blank\">https://www.kaggle.com/code/andrey67/uspto-annealing-lb-0-31</a></p>\n<p>Of course, I could not find MAGIC at all. It is very amazing who found it!</p>",
      "rawMarkdown": "Thank you for organizing this very interesting competition. And thank you to all participants.\nThe 22nd place solution is a simple solution using only “title” and “cpc”.\n\n# Creating candidates\n\n### Step 1: 90 candidates by TF-IDF (title)\n- stopwords='english'\n- ngram=(1, 3)\n- max_df=100\n- min_df=3\n- If a word is in another word set, exclude it\n- Exclude words with less than 5 letters\n\nStep 1 may not reach 60 because the title is too short. In that case, add Step 2 to the candidates.\n\n### Step 2: 60 candidates by TF-IDF (cpc)\n- max_df=0.01\n\nThe sum of candidates in Step 1 and Step 2 may not reach 30 because some publication_numbers have very little information on title and cpc. In that case, add Step 3 to the candidates.\n\n### Step 3: 10 candidates by word occurrence of title\n- Split title into words and sorted by number of occurrences (max 10)\n\n# Annealing algorithm\n\nUsing these candidates, run the Annealing algorithm. However, I am unable to understand the annealing algorithm and use it directly from the following two excellent public notebooks.\n@tubotubo @andrey67 I really appreciate it.\n\nhttps://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline\nhttps://www.kaggle.com/code/andrey67/uspto-annealing-lb-0-31\n\nOf course, I could not find MAGIC at all. It is very amazing who found it!",
      "votes": null
    },
    {
      "id": "2936358",
      "postDate": "07/26/2024 04:57:37",
      "content": "<p>Thank you for sharing your 22nd place solution! It's impressive how you utilized a simple yet effective TF-IDF approach with titles and CPC, and your step-by-step candidate creation process is very insightful. Great job! <a href=\"https://www.kaggle.com/thajime\" target=\"_blank\">@thajime</a> </p>",
      "rawMarkdown": "Thank you for sharing your 22nd place solution! It's impressive how you utilized a simple yet effective TF-IDF approach with titles and CPC, and your step-by-step candidate creation process is very insightful. Great job! @thajime",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2936358,
      "author_name": "aaditshukla",
      "author_url": "",
      "post_date": "07/26/2024 04:57:37",
      "content": "<p>Thank you for sharing your 22nd place solution! It's impressive how you utilized a simple yet effective TF-IDF approach with titles and CPC, and your step-by-step candidate creation process is very insightful. Great job! <a href=\"https://www.kaggle.com/thajime\" target=\"_blank\">@thajime</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2935831": "Thank you for organizing this very interesting competition. And thank you to all participants.\nThe 22nd place solution is a simple solution using only “title” and “cpc”.\n\n# Creating candidates\n\n### Step 1: 90 candidates by TF-IDF (title)\n- stopwords='english'\n- ngram=(1, 3)\n- max_df=100\n- min_df=3\n- If a word is in another word set, exclude it\n- Exclude words with less than 5 letters\n\nStep 1 may not reach 60 because the title is too short. In that case, add Step 2 to the candidates.\n\n### Step 2: 60 candidates by TF-IDF (cpc)\n- max_df=0.01\n\nThe sum of candidates in Step 1 and Step 2 may not reach 30 because some publication_numbers have very little information on title and cpc. In that case, add Step 3 to the candidates.\n\n### Step 3: 10 candidates by word occurrence of title\n- Split title into words and sorted by number of occurrences (max 10)\n\n# Annealing algorithm\n\nUsing these candidates, run the Annealing algorithm. However, I am unable to understand the annealing algorithm and use it directly from the following two excellent public notebooks.\n@tubotubo @andrey67 I really appreciate it.\n\nhttps://www.kaggle.com/code/tubotubo/uspto-simulated-annealing-baseline\nhttps://www.kaggle.com/code/andrey67/uspto-annealing-lb-0-31\n\nOf course, I could not find MAGIC at all. It is very amazing who found it!",
    "2936358": "Thank you for sharing your 22nd place solution! It's impressive how you utilized a simple yet effective TF-IDF approach with titles and CPC, and your step-by-step candidate creation process is very insightful. Great job! @thajime"
  },
  "source": "meta"
}