{
  "id": 509640,
  "title": "What does evaluation look like in details?",
  "url": "/competitions/uspto-explainable-ai/discussion/509640",
  "author_name": "Pavel Kazlou",
  "post_date": "2024-06-03T09:26:32.112000",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>After reading the problem statement and various discussions I still can't wrap my head around what data we can use when generating query.</p>\n<p>I can see test.csv file, which contains 50 nearest neighbours. (\"<em>A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</em>\").</p>\n<p>So the actual questions I have:</p>\n<ol>\n<li><p>When creating submission, do we need to generate queries for every patent found in 'publication_number' column of test.csv ?</p></li>\n<li><p>Am I right assuming, that when evaluating notebook, test.csv is replaced to contain 2,500 patents <strong>together with their top-50 neighbors</strong> ?</p></li>\n<li><p>If answer to previous is yes, then do those top-50 neighbors represent the ground truth for the purpose of mAP@50 metric? From the description: <em>Submissions are evaluated using mean average precision at 50 (mAP@50) between the patents your queries retrieve and the provided, related patent set.</em>  So are those top-50 neighbors represent \"provided, related patent set\"?  </p></li>\n<li><p>If answer to previous is yes, then I don't get the actual use case for our problem. From the description of use case, I would assume we need to generate query given just a single patent. So that the patent specialist can use that query to find other similar patents. But if instead patent specialist has to provide all the closest patents himselft in order to be able to generate query, then I don't get how our model can be useful. Do I miss something?</p></li>\n</ol>\n<p>I will highly appreciate any help in answering the questions.</p>",
  "messages": [
    {
      "id": 2852403,
      "postDate": "2024-06-03T09:26:32.113Z",
      "content": "<p>After reading the problem statement and various discussions I still can't wrap my head around what data we can use when generating query.</p>\n<p>I can see test.csv file, which contains 50 nearest neighbours. (\"<em>A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</em>\").</p>\n<p>So the actual questions I have:</p>\n<ol>\n<li><p>When creating submission, do we need to generate queries for every patent found in 'publication_number' column of test.csv ?</p></li>\n<li><p>Am I right assuming, that when evaluating notebook, test.csv is replaced to contain 2,500 patents <strong>together with their top-50 neighbors</strong> ?</p></li>\n<li><p>If answer to previous is yes, then do those top-50 neighbors represent the ground truth for the purpose of mAP@50 metric? From the description: <em>Submissions are evaluated using mean average precision at 50 (mAP@50) between the patents your queries retrieve and the provided, related patent set.</em>  So are those top-50 neighbors represent \"provided, related patent set\"?  </p></li>\n<li><p>If answer to previous is yes, then I don't get the actual use case for our problem. From the description of use case, I would assume we need to generate query given just a single patent. So that the patent specialist can use that query to find other similar patents. But if instead patent specialist has to provide all the closest patents himselft in order to be able to generate query, then I don't get how our model can be useful. Do I miss something?</p></li>\n</ol>\n<p>I will highly appreciate any help in answering the questions.</p>",
      "rawMarkdown": "After reading the problem statement and various discussions I still can't wrap my head around what data we can use when generating query.\n\nI can see test.csv file, which contains 50 nearest neighbours. (\"*A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.*\").\n\nSo the actual questions I have:\n\n1. When creating submission, do we need to generate queries for every patent found in 'publication_number' column of test.csv ?\n\n2. Am I right assuming, that when evaluating notebook, test.csv is replaced to contain 2,500 patents **together with their top-50 neighbors** ?\n\n3. If answer to previous is yes, then do those top-50 neighbors represent the ground truth for the purpose of mAP@50 metric? From the description: *Submissions are evaluated using mean average precision at 50 (mAP@50) between the patents your queries retrieve and the provided, related patent set.*  So are those top-50 neighbors represent \"provided, related patent set\"?  \n\n4. If answer to previous is yes, then I don't get the actual use case for our problem. From the description of use case, I would assume we need to generate query given just a single patent. So that the patent specialist can use that query to find other similar patents. But if instead patent specialist has to provide all the closest patents himselft in order to be able to generate query, then I don't get how our model can be useful. Do I miss something?\n\nI will highly appreciate any help in answering the questions.",
      "votes": 7
    },
    {
      "id": 2853097,
      "postDate": "2024-06-03T15:38:29.630Z",
      "content": "<ol>\n<li>Yes.</li>\n<li>Yes, you will get the top 50 neighbors in test.csv, but I am not sure whether there will be 2,500 rows or 2,500 patents in test.csv. I assume it is the former, 2,500 rows.</li>\n<li>Yes.</li>\n<li>Yes. I think the fundamental misunderstanding here is \"patent specialist has to provide all the closest patents himself\". If I understand the competition description correctly, there are AI models and other tools available to patent specialists that can automatically find patents that are similar to another, but that cannot explain why these patents are similar in a way that the patent specialist can understand. By finding a search query that matches a group of patents, one is able to explain why such a model or tool may have decided why certain patents are similar in a language that is familiar to these patent specialists.</li>\n</ol>",
      "rawMarkdown": "1. Yes.\n2. Yes, you will get the top 50 neighbors in test.csv, but I am not sure whether there will be 2,500 rows or 2,500 patents in test.csv. I assume it is the former, 2,500 rows.\n3. Yes.\n4. Yes. I think the fundamental misunderstanding here is \"patent specialist has to provide all the closest patents himself\". If I understand the competition description correctly, there are AI models and other tools available to patent specialists that can automatically find patents that are similar to another, but that cannot explain why these patents are similar in a way that the patent specialist can understand. By finding a search query that matches a group of patents, one is able to explain why such a model or tool may have decided why certain patents are similar in a language that is familiar to these patent specialists.",
      "votes": 8,
      "replies": [
        {
          "id": 2853191,
          "postDate": "2024-06-03T16:34:52.180Z",
          "content": "<p>That all looks right to me.</p>",
          "rawMarkdown": "That all looks right to me.",
          "votes": 1,
          "replies": [
            {
              "id": 2857677,
              "postDate": "2024-06-06T04:49:08.483Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Are u able to confirm if it's 50 rows or 2k5 rows on the actual test.csv?</p>",
              "rawMarkdown": "@sohier Are u able to confirm if it's 50 rows or 2k5 rows on the actual test.csv?",
              "votes": 1
            }
          ]
        },
        {
          "id": 2853288,
          "postDate": "2024-06-03T17:44:24.350Z",
          "content": "<p>Thanks, your explanation totally makes sense!</p>",
          "rawMarkdown": "Thanks, your explanation totally makes sense!"
        }
      ]
    },
    {
      "id": 2858333,
      "postDate": "2024-06-06T12:33:21.010Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2853097,
      "author_name": "Jasper",
      "author_url": "",
      "post_date": "2024-06-03T15:38:29.630000",
      "content": "<ol>\n<li>Yes.</li>\n<li>Yes, you will get the top 50 neighbors in test.csv, but I am not sure whether there will be 2,500 rows or 2,500 patents in test.csv. I assume it is the former, 2,500 rows.</li>\n<li>Yes.</li>\n<li>Yes. I think the fundamental misunderstanding here is \"patent specialist has to provide all the closest patents himself\". If I understand the competition description correctly, there are AI models and other tools available to patent specialists that can automatically find patents that are similar to another, but that cannot explain why these patents are similar in a way that the patent specialist can understand. By finding a search query that matches a group of patents, one is able to explain why such a model or tool may have decided why certain patents are similar in a language that is familiar to these patent specialists.</li>\n</ol>",
      "votes": 8,
      "replies": [
        {
          "id": 2853191,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2024-06-03T16:34:52.180000",
          "content": "<p>That all looks right to me.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2857677,
              "author_name": "DH",
              "author_url": "",
              "post_date": "2024-06-06T04:49:08.483000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Are u able to confirm if it's 50 rows or 2k5 rows on the actual test.csv?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2853288,
          "author_name": "Pavel Kazlou",
          "author_url": "",
          "post_date": "2024-06-03T17:44:24.350000",
          "content": "<p>Thanks, your explanation totally makes sense!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2858333,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-06T12:33:21.010000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2852403": "After reading the problem statement and various discussions I still can't wrap my head around what data we can use when generating query.\n\nI can see test.csv file, which contains 50 nearest neighbours. (\"*A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.*\").\n\nSo the actual questions I have:\n\n1. When creating submission, do we need to generate queries for every patent found in 'publication_number' column of test.csv ?\n\n2. Am I right assuming, that when evaluating notebook, test.csv is replaced to contain 2,500 patents **together with their top-50 neighbors** ?\n\n3. If answer to previous is yes, then do those top-50 neighbors represent the ground truth for the purpose of mAP@50 metric? From the description: *Submissions are evaluated using mean average precision at 50 (mAP@50) between the patents your queries retrieve and the provided, related patent set.*  So are those top-50 neighbors represent \"provided, related patent set\"?  \n\n4. If answer to previous is yes, then I don't get the actual use case for our problem. From the description of use case, I would assume we need to generate query given just a single patent. So that the patent specialist can use that query to find other similar patents. But if instead patent specialist has to provide all the closest patents himselft in order to be able to generate query, then I don't get how our model can be useful. Do I miss something?\n\nI will highly appreciate any help in answering the questions.",
    "2853097": "1. Yes.\n2. Yes, you will get the top 50 neighbors in test.csv, but I am not sure whether there will be 2,500 rows or 2,500 patents in test.csv. I assume it is the former, 2,500 rows.\n3. Yes.\n4. Yes. I think the fundamental misunderstanding here is \"patent specialist has to provide all the closest patents himself\". If I understand the competition description correctly, there are AI models and other tools available to patent specialists that can automatically find patents that are similar to another, but that cannot explain why these patents are similar in a way that the patent specialist can understand. By finding a search query that matches a group of patents, one is able to explain why such a model or tool may have decided why certain patents are similar in a language that is familiar to these patent specialists.",
    "2858333": ""
  }
}