{
  "id": 317437,
  "title": "Metric Evaluation Code",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/317437",
  "author_name": "",
  "post_date": "2022-04-07T04:53:02.056084Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>To calculate the metric score for my prediction i am using the following code:-</p>\n<pre><code>def apk(actual, predicted, k=10):\n    if len(predicted)&gt;k:\n        predicted = predicted[:k]\n\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=10):\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n</code></pre>\n<p>My submission on using most popular 12 items score .003 but my cv for last week comes .0011.Is it because there is something wrong in my metric code?</p>",
  "messages": [
    {
      "id": "1747837",
      "postDate": "04/07/2022 04:53:02",
      "content": "<p>To calculate the metric score for my prediction i am using the following code:-</p>\n<pre><code>def apk(actual, predicted, k=10):\n    if len(predicted)&gt;k:\n        predicted = predicted[:k]\n\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=10):\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n</code></pre>\n<p>My submission on using most popular 12 items score .003 but my cv for last week comes .0011.Is it because there is something wrong in my metric code?</p>",
      "rawMarkdown": "To calculate the metric score for my prediction i am using the following code:-\n```\ndef apk(actual, predicted, k=10):\n    if len(predicted)>k:\n        predicted = predicted[:k]\n\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=10):\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n```\nMy submission on using most popular 12 items score .003 but my cv for last week comes .0011.Is it because there is something wrong in my metric code?",
      "votes": null
    },
    {
      "id": "1748386",
      "postDate": "04/07/2022 14:40:35",
      "content": "<p>Not necessarily so - there does seem to be a difference between cv and LB.</p>\n<p>I haven't run that exact experiment, but have also come across similar results.</p>",
      "rawMarkdown": "Not necessarily so - there does seem to be a difference between cv and LB.\n\nI haven't run that exact experiment, but have also come across similar results.",
      "votes": null
    },
    {
      "id": "1748391",
      "postDate": "04/07/2022 14:42:46",
      "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> So the metric evaluation is correct ?</p>",
      "rawMarkdown": "jacob34 So the metric evaluation is correct ?",
      "votes": null
    },
    {
      "id": "1748427",
      "postDate": "04/07/2022 15:22:08",
      "content": "<p>Reading through it now…</p>\n<p>default value for k should be 12, not 10.</p>\n<p>In mapk:<br>\n<code>tmp = predicted.merge(actual,on = 'customer_id')</code></p>\n<p>that's an inner join, so will only return customers present in both dfs.</p>\n<ol>\n<li>To get accurate results, you need <code>actual</code> to only have customers who bought something - otherwise you could evaluate on customers who never bought anything.</li>\n<li>You also need <code>predicted</code> to not be missing any customers who bought something - or else you won't be evaluating on them.</li>\n</ol>\n<p>For #2, you can do:<br>\n<code>tmp = actual.merge(predicted, on='customer_id', how='left')</code></p>",
      "rawMarkdown": "Reading through it now...\n\ndefault value for k should be 12, not 10.\n\nIn mapk:\n`tmp = predicted.merge(actual,on = 'customer_id')`\n\nthat's an inner join, so will only return customers present in both dfs.\n1. To get accurate results, you need `actual` to only have customers who bought something - otherwise you could evaluate on customers who never bought anything.\n2. You also need `predicted` to not be missing any customers who bought something - or else you won't be evaluating on them.\n\nFor #2, you can do:\n`tmp = actual.merge(predicted, on='customer_id', how='left')`",
      "votes": null
    },
    {
      "id": "1748440",
      "postDate": "04/07/2022 15:50:15",
      "content": "<p>I added some assertion statements to be sure I don't miss such cases.I guess it's working fine now.</p>\n<pre><code>def apk(actual, predicted, k=12):\n    if len(predicted)&gt;k:\n        predicted = predicted[:k]\n    if len(actual)==0:\n      assert False\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=12):\n    assert 1371980 == len(predicted) # Number of Customers in Sample sub\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n</code></pre>",
      "rawMarkdown": "I added some assertion statements to be sure I don't miss such cases.I guess it's working fine now.\n```\ndef apk(actual, predicted, k=12):\n    if len(predicted)>k:\n        predicted = predicted[:k]\n    if len(actual)==0:\n      assert False\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=12):\n    assert 1371980 == len(predicted) # Number of Customers in Sample sub\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1748386,
      "author_name": "jacob34",
      "author_url": "",
      "post_date": "04/07/2022 14:40:35",
      "content": "<p>Not necessarily so - there does seem to be a difference between cv and LB.</p>\n<p>I haven't run that exact experiment, but have also come across similar results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1748391,
          "author_name": "devanshchowdhury",
          "author_url": "",
          "post_date": "04/07/2022 14:42:46",
          "content": "<p><a href=\"https://www.kaggle.com/jacob34\" target=\"_blank\">@jacob34</a> So the metric evaluation is correct ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1748427,
          "author_name": "jacob34",
          "author_url": "",
          "post_date": "04/07/2022 15:22:08",
          "content": "<p>Reading through it now…</p>\n<p>default value for k should be 12, not 10.</p>\n<p>In mapk:<br>\n<code>tmp = predicted.merge(actual,on = 'customer_id')</code></p>\n<p>that's an inner join, so will only return customers present in both dfs.</p>\n<ol>\n<li>To get accurate results, you need <code>actual</code> to only have customers who bought something - otherwise you could evaluate on customers who never bought anything.</li>\n<li>You also need <code>predicted</code> to not be missing any customers who bought something - or else you won't be evaluating on them.</li>\n</ol>\n<p>For #2, you can do:<br>\n<code>tmp = actual.merge(predicted, on='customer_id', how='left')</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1748440,
          "author_name": "devanshchowdhury",
          "author_url": "",
          "post_date": "04/07/2022 15:50:15",
          "content": "<p>I added some assertion statements to be sure I don't miss such cases.I guess it's working fine now.</p>\n<pre><code>def apk(actual, predicted, k=12):\n    if len(predicted)&gt;k:\n        predicted = predicted[:k]\n    if len(actual)==0:\n      assert False\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=12):\n    assert 1371980 == len(predicted) # Number of Customers in Sample sub\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1747837": "To calculate the metric score for my prediction i am using the following code:-\n```\ndef apk(actual, predicted, k=10):\n    if len(predicted)>k:\n        predicted = predicted[:k]\n\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=10):\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n```\nMy submission on using most popular 12 items score .003 but my cv for last week comes .0011.Is it because there is something wrong in my metric code?",
    "1748386": "Not necessarily so - there does seem to be a difference between cv and LB.\n\nI haven't run that exact experiment, but have also come across similar results.",
    "1748391": "jacob34 So the metric evaluation is correct ?",
    "1748427": "Reading through it now...\n\ndefault value for k should be 12, not 10.\n\nIn mapk:\n`tmp = predicted.merge(actual,on = 'customer_id')`\n\nthat's an inner join, so will only return customers present in both dfs.\n1. To get accurate results, you need `actual` to only have customers who bought something - otherwise you could evaluate on customers who never bought anything.\n2. You also need `predicted` to not be missing any customers who bought something - or else you won't be evaluating on them.\n\nFor #2, you can do:\n`tmp = actual.merge(predicted, on='customer_id', how='left')`",
    "1748440": "I added some assertion statements to be sure I don't miss such cases.I guess it's working fine now.\n```\ndef apk(actual, predicted, k=12):\n    if len(predicted)>k:\n        predicted = predicted[:k]\n    if len(actual)==0:\n      assert False\n    score = 0.0\n    num_hits = 0.0\n    predicted = predicted.split()\n    actual = actual.split()\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=12):\n    assert 1371980 == len(predicted) # Number of Customers in Sample sub\n    tmp = predicted.merge(actual,on = 'customer_id')\n    return np.mean([apk(a,p,k) for a,p in zip(tmp.article_id, tmp.prediction)])\n```"
  },
  "source": "meta"
}