{
  "id": 319327,
  "title": "How to evaluate generated candidates?",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/319327",
  "author_name": "",
  "post_date": "2022-04-16T17:04:41.367665500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The purpose of the competition is to recommend 12 items. In general, it is necessary to generate candidates to be recommended before the re-ranking step.</p>\n<p>However, it is very difficult to properly generate an personalize candidate group, and it is considered to be the main task of the competition. Therefore, it is also important to evaluate the generated candidates through various techniques.</p>\n<p>I am evaluating candidates based on the following three indicators.</p>\n<ul>\n<li>Actual Article_ids</li>\n</ul>\n<pre><code>test = transactions.query('week==0').copy().reset_index(drop=True)\ntest_actual = test.groupby('customer_id')['article_id'].apply(list)\n</code></pre>\n<ul>\n<li>generated candidates</li>\n</ul>\n<pre><code>candi = generate_candidates(test['customer_id'].unique(), n=k)\ntest_pred = (candi.groupby('customer_id')['article_id'].apply(list))\n</code></pre>\n<p><strong>1. Inner Ratio</strong></p>\n<pre><code>actual = set(test.article_id.values)\npred = set(candi.article_id.values)\n\nprint(1 - (len(actual-pred) / len(actual)))\n</code></pre>\n<p>If the ratio is low, it can be interpreted that the variety of items to be recommended decreases. Furthermore, it seems that it can be measured with an index such as entropy diversity.</p>\n<p><strong>2. Hits Ratio</strong></p>\n<pre><code>def hit(actual_item, pred_items):\n    hits = 0\n    for i in actual_item:\n        if i in pred_items:\n            hits+=1\n    hits_ratio = hits / len(actual_item)\n    return hits_ratio\n\ndef hit_at_generated(actual, pred):\n    return np.mean([hit(a, p) for a, p in zip(actual, pred)])\n</code></pre>\n<p><code>hit_score = hit_at_generated(test_actual.values, test_pred.values)</code></p>\n<p>The purpose of the competition is to increase the map@12 score, but in the generate candidation stage, the priority is to match the items to be recommended.</p>\n<p><strong>3. Size</strong><br>\n<code>print(candi.shape[0])</code></p>\n<p>In my opinion, the optimal size is 12 items per customer.</p>\n<p>According to the above three indicators, my current score is as follows.<br>\n[0.51, 0.1642, 36413992]</p>\n<p>Please share your score and let me know if there is a better way to evaluate candidates.</p>",
  "messages": [
    {
      "id": "1757475",
      "postDate": "04/16/2022 17:04:41",
      "content": "<p>The purpose of the competition is to recommend 12 items. In general, it is necessary to generate candidates to be recommended before the re-ranking step.</p>\n<p>However, it is very difficult to properly generate an personalize candidate group, and it is considered to be the main task of the competition. Therefore, it is also important to evaluate the generated candidates through various techniques.</p>\n<p>I am evaluating candidates based on the following three indicators.</p>\n<ul>\n<li>Actual Article_ids</li>\n</ul>\n<pre><code>test = transactions.query('week==0').copy().reset_index(drop=True)\ntest_actual = test.groupby('customer_id')['article_id'].apply(list)\n</code></pre>\n<ul>\n<li>generated candidates</li>\n</ul>\n<pre><code>candi = generate_candidates(test['customer_id'].unique(), n=k)\ntest_pred = (candi.groupby('customer_id')['article_id'].apply(list))\n</code></pre>\n<p><strong>1. Inner Ratio</strong></p>\n<pre><code>actual = set(test.article_id.values)\npred = set(candi.article_id.values)\n\nprint(1 - (len(actual-pred) / len(actual)))\n</code></pre>\n<p>If the ratio is low, it can be interpreted that the variety of items to be recommended decreases. Furthermore, it seems that it can be measured with an index such as entropy diversity.</p>\n<p><strong>2. Hits Ratio</strong></p>\n<pre><code>def hit(actual_item, pred_items):\n    hits = 0\n    for i in actual_item:\n        if i in pred_items:\n            hits+=1\n    hits_ratio = hits / len(actual_item)\n    return hits_ratio\n\ndef hit_at_generated(actual, pred):\n    return np.mean([hit(a, p) for a, p in zip(actual, pred)])\n</code></pre>\n<p><code>hit_score = hit_at_generated(test_actual.values, test_pred.values)</code></p>\n<p>The purpose of the competition is to increase the map@12 score, but in the generate candidation stage, the priority is to match the items to be recommended.</p>\n<p><strong>3. Size</strong><br>\n<code>print(candi.shape[0])</code></p>\n<p>In my opinion, the optimal size is 12 items per customer.</p>\n<p>According to the above three indicators, my current score is as follows.<br>\n[0.51, 0.1642, 36413992]</p>\n<p>Please share your score and let me know if there is a better way to evaluate candidates.</p>",
      "rawMarkdown": "The purpose of the competition is to recommend 12 items. In general, it is necessary to generate candidates to be recommended before the re-ranking step.\n\nHowever, it is very difficult to properly generate an personalize candidate group, and it is considered to be the main task of the competition. Therefore, it is also important to evaluate the generated candidates through various techniques.\n\nI am evaluating candidates based on the following three indicators.\n\n- Actual Article_ids\n```\ntest = transactions.query('week==0').copy().reset_index(drop=True)\ntest_actual = test.groupby('customer_id')['article_id'].apply(list)\n```\n\n- generated candidates\n```\ncandi = generate_candidates(test['customer_id'].unique(), n=k)\ntest_pred = (candi.groupby('customer_id')['article_id'].apply(list))\n```\n\n**1. Inner Ratio**\n\n```\nactual = set(test.article_id.values)\npred = set(candi.article_id.values)\n\nprint(1 - (len(actual-pred) / len(actual)))\n```\nIf the ratio is low, it can be interpreted that the variety of items to be recommended decreases. Furthermore, it seems that it can be measured with an index such as entropy diversity.\n\n**2. Hits Ratio**\n\n```\ndef hit(actual_item, pred_items):\n    hits = 0\n    for i in actual_item:\n        if i in pred_items:\n            hits+=1\n    hits_ratio = hits / len(actual_item)\n    return hits_ratio\n\ndef hit_at_generated(actual, pred):\n    return np.mean([hit(a, p) for a, p in zip(actual, pred)])\n```\n\n`hit_score = hit_at_generated(test_actual.values, test_pred.values)`\n\nThe purpose of the competition is to increase the map@12 score, but in the generate candidation stage, the priority is to match the items to be recommended.\n\n**3. Size**\n```print(candi.shape[0])```\n\nIn my opinion, the optimal size is 12 items per customer.\n\n\nAccording to the above three indicators, my current score is as follows.\n[0.51, 0.1642, 36413992]\n\nPlease share your score and let me know if there is a better way to evaluate candidates.",
      "votes": null
    },
    {
      "id": "1757816",
      "postDate": "04/17/2022 03:40:37",
      "content": "<p>I calculate the Maximum of MAP 12 from the candidates in Last Week CV</p>\n<p>more detail on this <a href=\"https://www.kaggle.com/code/hervind/h-m-how-to-evaluate-candidate/notebook\" target=\"_blank\">https://www.kaggle.com/code/hervind/h-m-how-to-evaluate-candidate/notebook</a></p>",
      "rawMarkdown": "I calculate the Maximum of MAP 12 from the candidates in Last Week CV\n\nmore detail on this https://www.kaggle.com/code/hervind/h-m-how-to-evaluate-candidate/notebook",
      "votes": null
    },
    {
      "id": "1758211",
      "postDate": "04/17/2022 12:46:14",
      "content": "<p>Thanks for sharing.<br>\nIn your notebook, covered customer indicator is very useful!</p>",
      "rawMarkdown": "Thanks for sharing.\nIn your notebook, covered customer indicator is very useful!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1757816,
      "author_name": "hervind",
      "author_url": "",
      "post_date": "04/17/2022 03:40:37",
      "content": "<p>I calculate the Maximum of MAP 12 from the candidates in Last Week CV</p>\n<p>more detail on this <a href=\"https://www.kaggle.com/code/hervind/h-m-how-to-evaluate-candidate/notebook\" target=\"_blank\">https://www.kaggle.com/code/hervind/h-m-how-to-evaluate-candidate/notebook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1758211,
          "author_name": "pyy0715",
          "author_url": "",
          "post_date": "04/17/2022 12:46:14",
          "content": "<p>Thanks for sharing.<br>\nIn your notebook, covered customer indicator is very useful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1757475": "The purpose of the competition is to recommend 12 items. In general, it is necessary to generate candidates to be recommended before the re-ranking step.\n\nHowever, it is very difficult to properly generate an personalize candidate group, and it is considered to be the main task of the competition. Therefore, it is also important to evaluate the generated candidates through various techniques.\n\nI am evaluating candidates based on the following three indicators.\n\n- Actual Article_ids\n```\ntest = transactions.query('week==0').copy().reset_index(drop=True)\ntest_actual = test.groupby('customer_id')['article_id'].apply(list)\n```\n\n- generated candidates\n```\ncandi = generate_candidates(test['customer_id'].unique(), n=k)\ntest_pred = (candi.groupby('customer_id')['article_id'].apply(list))\n```\n\n**1. Inner Ratio**\n\n```\nactual = set(test.article_id.values)\npred = set(candi.article_id.values)\n\nprint(1 - (len(actual-pred) / len(actual)))\n```\nIf the ratio is low, it can be interpreted that the variety of items to be recommended decreases. Furthermore, it seems that it can be measured with an index such as entropy diversity.\n\n**2. Hits Ratio**\n\n```\ndef hit(actual_item, pred_items):\n    hits = 0\n    for i in actual_item:\n        if i in pred_items:\n            hits+=1\n    hits_ratio = hits / len(actual_item)\n    return hits_ratio\n\ndef hit_at_generated(actual, pred):\n    return np.mean([hit(a, p) for a, p in zip(actual, pred)])\n```\n\n`hit_score = hit_at_generated(test_actual.values, test_pred.values)`\n\nThe purpose of the competition is to increase the map@12 score, but in the generate candidation stage, the priority is to match the items to be recommended.\n\n**3. Size**\n```print(candi.shape[0])```\n\nIn my opinion, the optimal size is 12 items per customer.\n\n\nAccording to the above three indicators, my current score is as follows.\n[0.51, 0.1642, 36413992]\n\nPlease share your score and let me know if there is a better way to evaluate candidates.",
    "1757816": "I calculate the Maximum of MAP 12 from the candidates in Last Week CV\n\nmore detail on this https://www.kaggle.com/code/hervind/h-m-how-to-evaluate-candidate/notebook",
    "1758211": "Thanks for sharing.\nIn your notebook, covered customer indicator is very useful!"
  },
  "source": "meta"
}