{
  "id": 499981,
  "title": "mAP without probabilities?",
  "url": "/competitions/uspto-explainable-ai/discussion/499981",
  "author_name": "",
  "post_date": "2024-05-03T18:10:48.730466900Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have difficulty understanding the competition metric. How can we calculate mAP without scores or probabilities to sort with? Boolean search doesn't have any ordering between its outputs. </p>",
  "messages": [
    {
      "id": "2791563",
      "postDate": "05/03/2024 18:10:48",
      "content": "<p>I have difficulty understanding the competition metric. How can we calculate mAP without scores or probabilities to sort with? Boolean search doesn't have any ordering between its outputs. </p>",
      "rawMarkdown": "I have difficulty understanding the competition metric. How can we calculate mAP without scores or probabilities to sort with? Boolean search doesn't have any ordering between its outputs.",
      "votes": null
    },
    {
      "id": "2791642",
      "postDate": "05/03/2024 19:08:41",
      "content": "<p>If you do a search for \"mean average precision information retrieval\" you will plenty of articles explaining it. Here is one: <a href=\"https://towardsdatascience.com/breaking-down-mean-average-precision-map-ae462f623a52\" target=\"_blank\">https://towardsdatascience.com/breaking-down-mean-average-precision-map-ae462f623a52</a></p>\n<p>I actually think the competition metric is incorrectly implemented(at least according to all the definitions I have seen). </p>\n<p>I believe the competition implements average precision like this (val score: 0.63ish, lb:62):</p>\n<pre><code> ():\n    precisions = ()\n    n_label = (labels)\n    n_found = \n     e, i  (preds):\n         i  labels:\n            n_found += \n        precisions.append(n_found/(e+)) \n     (precisions)/\n</code></pre>\n<p>When all the articles say it should be implemented like this (val score 0.32 or so, lb:62)</p>\n<pre><code> ():\n    precisions = ()\n    n_label = (labels)\n    n_found = \n     e, i  (preds):\n         i  labels:\n            n_found += \n            precisions.append(n_found/(e+)) \n     (precisions)/\n</code></pre>\n<p>I didn't know there was no order to the search results(naively thought <code>a OR b</code> would return results match <code>a</code> first but now that I think about it that doesn't have to be) . That does make it difficult to control false positives.</p>",
      "rawMarkdown": "If you do a search for \"mean average precision information retrieval\" you will plenty of articles explaining it. Here is one: https://towardsdatascience.com/breaking-down-mean-average-precision-map-ae462f623a52\n\nI actually think the competition metric is incorrectly implemented(at least according to all the definitions I have seen). \n\nI believe the competition implements average precision like this (val score: 0.63ish, lb:62):\n```\ndef ap50(preds, labels):\n    precisions = list()\n    n_label = len(labels)\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n        precisions.append(n_found/(e+1)) # this is the line that is probably incorrect for competition \n    return sum(precisions)/50\n```\n\nWhen all the articles say it should be implemented like this (val score 0.32 or so, lb:62)\n```\ndef ap50(preds, labels):\n    precisions = list()\n    n_label = len(labels)\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n            precisions.append(n_found/(e+1)) # this is how it probably should be\n    return sum(precisions)/50\n```\n\nI didn't know there was no order to the search results(naively thought `a OR b` would return results match `a` first but now that I think about it that doesn't have to be) . That does make it difficult to control false positives.",
      "votes": null
    },
    {
      "id": "2792387",
      "postDate": "05/04/2024 07:41:36",
      "content": "<p>From what I have experimented with, specifying ´terms=True´ in a search as explained <a href=\"https://whoosh.readthedocs.io/en/latest/searching.html#which-terms-from-my-query-matched\" target=\"_blank\">here</a> actually would sort results by how many items matched - however, I think we do not know how this is implemented during evaluation.</p>",
      "rawMarkdown": "From what I have experimented with, specifying ´terms=True´ in a search as explained [here](https://whoosh.readthedocs.io/en/latest/searching.html#which-terms-from-my-query-matched) actually would sort results by how many items matched - however, I think we do not know how this is implemented during evaluation.",
      "votes": null
    },
    {
      "id": "2881639",
      "postDate": "06/20/2024 21:27:26",
      "content": "<p>I found a pretty strong match between validation on F1 and LB score so far. That symmetry could break later on but for now I'm just using it as a proxy. </p>",
      "rawMarkdown": "I found a pretty strong match between validation on F1 and LB score so far. That symmetry could break later on but for now I'm just using it as a proxy.",
      "votes": null
    },
    {
      "id": "2890420",
      "postDate": "06/26/2024 05:49:33",
      "content": "<p>what did you mean by validation on F1? also trying to get a proxy for LB scores</p>",
      "rawMarkdown": "what did you mean by validation on F1? also trying to get a proxy for LB scores",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2791642,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "05/03/2024 19:08:41",
      "content": "<p>If you do a search for \"mean average precision information retrieval\" you will plenty of articles explaining it. Here is one: <a href=\"https://towardsdatascience.com/breaking-down-mean-average-precision-map-ae462f623a52\" target=\"_blank\">https://towardsdatascience.com/breaking-down-mean-average-precision-map-ae462f623a52</a></p>\n<p>I actually think the competition metric is incorrectly implemented(at least according to all the definitions I have seen). </p>\n<p>I believe the competition implements average precision like this (val score: 0.63ish, lb:62):</p>\n<pre><code> ():\n    precisions = ()\n    n_label = (labels)\n    n_found = \n     e, i  (preds):\n         i  labels:\n            n_found += \n        precisions.append(n_found/(e+)) \n     (precisions)/\n</code></pre>\n<p>When all the articles say it should be implemented like this (val score 0.32 or so, lb:62)</p>\n<pre><code> ():\n    precisions = ()\n    n_label = (labels)\n    n_found = \n     e, i  (preds):\n         i  labels:\n            n_found += \n            precisions.append(n_found/(e+)) \n     (precisions)/\n</code></pre>\n<p>I didn't know there was no order to the search results(naively thought <code>a OR b</code> would return results match <code>a</code> first but now that I think about it that doesn't have to be) . That does make it difficult to control false positives.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2792387,
          "author_name": "valentinwerner",
          "author_url": "",
          "post_date": "05/04/2024 07:41:36",
          "content": "<p>From what I have experimented with, specifying ´terms=True´ in a search as explained <a href=\"https://whoosh.readthedocs.io/en/latest/searching.html#which-terms-from-my-query-matched\" target=\"_blank\">here</a> actually would sort results by how many items matched - however, I think we do not know how this is implemented during evaluation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2881639,
      "author_name": "raki21",
      "author_url": "",
      "post_date": "06/20/2024 21:27:26",
      "content": "<p>I found a pretty strong match between validation on F1 and LB score so far. That symmetry could break later on but for now I'm just using it as a proxy. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2890420,
          "author_name": "thedilpreet",
          "author_url": "",
          "post_date": "06/26/2024 05:49:33",
          "content": "<p>what did you mean by validation on F1? also trying to get a proxy for LB scores</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2791563": "I have difficulty understanding the competition metric. How can we calculate mAP without scores or probabilities to sort with? Boolean search doesn't have any ordering between its outputs.",
    "2791642": "If you do a search for \"mean average precision information retrieval\" you will plenty of articles explaining it. Here is one: https://towardsdatascience.com/breaking-down-mean-average-precision-map-ae462f623a52\n\nI actually think the competition metric is incorrectly implemented(at least according to all the definitions I have seen). \n\nI believe the competition implements average precision like this (val score: 0.63ish, lb:62):\n```\ndef ap50(preds, labels):\n    precisions = list()\n    n_label = len(labels)\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n        precisions.append(n_found/(e+1)) # this is the line that is probably incorrect for competition \n    return sum(precisions)/50\n```\n\nWhen all the articles say it should be implemented like this (val score 0.32 or so, lb:62)\n```\ndef ap50(preds, labels):\n    precisions = list()\n    n_label = len(labels)\n    n_found = 0\n    for e, i in enumerate(preds):\n        if i in labels:\n            n_found += 1\n            precisions.append(n_found/(e+1)) # this is how it probably should be\n    return sum(precisions)/50\n```\n\nI didn't know there was no order to the search results(naively thought `a OR b` would return results match `a` first but now that I think about it that doesn't have to be) . That does make it difficult to control false positives.",
    "2792387": "From what I have experimented with, specifying ´terms=True´ in a search as explained [here](https://whoosh.readthedocs.io/en/latest/searching.html#which-terms-from-my-query-matched) actually would sort results by how many items matched - however, I think we do not know how this is implemented during evaluation.",
    "2881639": "I found a pretty strong match between validation on F1 and LB score so far. That symmetry could break later on but for now I'm just using it as a proxy.",
    "2890420": "what did you mean by validation on F1? also trying to get a proxy for LB scores"
  },
  "source": "meta"
}