{
  "id": 505768,
  "title": "What is the difference between train_index and test_index?",
  "url": "/competitions/uspto-explainable-ai/discussion/505768",
  "author_name": "penguin46",
  "post_date": "2024-05-19T03:43:25.592000",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>The dataset description states the following. What exactly does \"equivalent in size and setup\" mean? </p>\n<blockquote>\n  <p>train_index_patent_ids.json A list of the patents included in the Whoosh index.<br>\n  train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975.</p>\n</blockquote>\n<p>I have the following questions:</p>\n<ol>\n<li><p>Is the size of test_index the same as train_index, 200,000?<br>\nThat is, does the whoosh index that is run during scoring contain the same number of patents as the train_index (exactly 200,000)?</p></li>\n<li><p>Are all the patents in test.csv included in test_index?<br>\nIn other words, is it possible to achieve a score close to 1.00 if ideal queries exist? This does not seem to be achievable in train_index because no matter how train.csv is constructed, it does not contain enough neighborhoods as descrived <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497956\" target=\"_blank\">here</a>.</p></li>\n<li><p>Does test_index also include the patents prior to 1975?<br>\nThe neighborhoods of the patent in train_index contain pre-1975 patents, which may violate the setup that \"Only includes patents published on or after 1975\". Can we assume that all test.csv neighbors are on or after 1975? Or are the test.csv neighborhoods included in test_index even if they are before 1975?</p></li>\n</ol>",
  "messages": [
    {
      "id": 2823099,
      "postDate": "2024-05-19T03:43:25.593Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>The dataset description states the following. What exactly does \"equivalent in size and setup\" mean? </p>\n<blockquote>\n  <p>train_index_patent_ids.json A list of the patents included in the Whoosh index.<br>\n  train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975.</p>\n</blockquote>\n<p>I have the following questions:</p>\n<ol>\n<li><p>Is the size of test_index the same as train_index, 200,000?<br>\nThat is, does the whoosh index that is run during scoring contain the same number of patents as the train_index (exactly 200,000)?</p></li>\n<li><p>Are all the patents in test.csv included in test_index?<br>\nIn other words, is it possible to achieve a score close to 1.00 if ideal queries exist? This does not seem to be achievable in train_index because no matter how train.csv is constructed, it does not contain enough neighborhoods as descrived <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497956\" target=\"_blank\">here</a>.</p></li>\n<li><p>Does test_index also include the patents prior to 1975?<br>\nThe neighborhoods of the patent in train_index contain pre-1975 patents, which may violate the setup that \"Only includes patents published on or after 1975\". Can we assume that all test.csv neighbors are on or after 1975? Or are the test.csv neighborhoods included in test_index even if they are before 1975?</p></li>\n</ol>",
      "rawMarkdown": "Hi @addisonhoward @sohier,\n\nThe dataset description states the following. What exactly does \"equivalent in size and setup\" mean? \n\n> train_index_patent_ids.json A list of the patents included in the Whoosh index.\n> train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975.\n\nI have the following questions:\n\n1. Is the size of test_index the same as train_index, 200,000?\nThat is, does the whoosh index that is run during scoring contain the same number of patents as the train_index (exactly 200,000)?\n\n2. Are all the patents in test.csv included in test_index?\nIn other words, is it possible to achieve a score close to 1.00 if ideal queries exist? This does not seem to be achievable in train_index because no matter how train.csv is constructed, it does not contain enough neighborhoods as descrived [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497956).\n\n3. Does test_index also include the patents prior to 1975?\nThe neighborhoods of the patent in train_index contain pre-1975 patents, which may violate the setup that \"Only includes patents published on or after 1975\". Can we assume that all test.csv neighbors are on or after 1975? Or are the test.csv neighborhoods included in test_index even if they are before 1975?",
      "votes": 18
    },
    {
      "id": 2853209,
      "postDate": "2024-06-03T16:52:11.163Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>Could you please answer these questions?<br>\nSorry if you are preparing answers, but I was wondering if perhaps you missed this thread.</p>",
      "rawMarkdown": "Hi @addisonhoward @sohier,\n\nCould you please answer these questions?\nSorry if you are preparing answers, but I was wondering if perhaps you missed this thread.",
      "votes": 1
    },
    {
      "id": 2823191,
      "postDate": "2024-05-19T04:39:35.117Z",
      "content": "<p>I've been thinking about exactly the same questions. I'm not one of the organizers, but my guess is as follows:</p>\n<ol>\n<li>It is likely to be about 200,000 but it isn't necessarily exactly 200,000.</li>\n<li>Yes, I'm pretty sure that all neighbors are included into the test index.</li>\n<li>Since the train index does include them, we have to assume that the test index contains them as well. Which is weird given all the wording about 1975.</li>\n</ol>\n<p>Also, speaking of achieving 1.0 score: as discussed <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981\" target=\"_blank\">here</a>, and I agree with it, the current mAP metric is being incorrectly calculated. So I won't be surprised if it gets fixed one day. The current scores will roughly be halved and 1.0 will be far away.</p>",
      "rawMarkdown": "I've been thinking about exactly the same questions. I'm not one of the organizers, but my guess is as follows:\n\n1. It is likely to be about 200,000 but it isn't necessarily exactly 200,000.\n2. Yes, I'm pretty sure that all neighbors are included into the test index.\n3. Since the train index does include them, we have to assume that the test index contains them as well. Which is weird given all the wording about 1975.\n\nAlso, speaking of achieving 1.0 score: as discussed [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981), and I agree with it, the current mAP metric is being incorrectly calculated. So I won't be surprised if it gets fixed one day. The current scores will roughly be halved and 1.0 will be far away.\n",
      "votes": 1,
      "replies": [
        {
          "id": 2823220,
          "postDate": "2024-05-19T04:57:44.300Z",
          "content": "<p>Thanks for the reply. I generally agree with you.<br>\nRegarding the size of test_index, I am concerned that it may contain up to an additional 125,000 (2500x50) to include all the neighborhoods in test.csv, in addition to the 200,000 that result from the same generation method as train_index.</p>\n<p>Let's wait for a reply from the host.</p>",
          "rawMarkdown": "Thanks for the reply. I generally agree with you.\nRegarding the size of test_index, I am concerned that it may contain up to an additional 125,000 (2500x50) to include all the neighborhoods in test.csv, in addition to the 200,000 that result from the same generation method as train_index.\n\nLet's wait for a reply from the host.",
          "replies": [
            {
              "id": 2853226,
              "postDate": "2024-06-03T17:07:35.250Z",
              "content": "<p>The test index contains approximately 200,000 rows. </p>",
              "rawMarkdown": "The test index contains approximately 200,000 rows. ",
              "votes": 4
            },
            {
              "id": 2853231,
              "postDate": "2024-06-03T17:12:38.903Z",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThank you for your reply!<br>\nCould you also answer the second and third questions?</p>",
              "rawMarkdown": "@sohier \nThank you for your reply!\nCould you also answer the second and third questions?",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2888328,
      "postDate": "2024-06-24T18:00:25.563Z",
      "content": "<p>i thought the train index has way more entries than 200,000? shouldnt it contain len(meta_data) * 50  of patents?</p>",
      "rawMarkdown": "i thought the train index has way more entries than 200,000? shouldnt it contain len(meta_data) * 50  of patents?"
    },
    {
      "id": 2858345,
      "postDate": "2024-06-06T12:40:19.340Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2853209,
      "author_name": "penguin46",
      "author_url": "",
      "post_date": "2024-06-03T16:52:11.163000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>Could you please answer these questions?<br>\nSorry if you are preparing answers, but I was wondering if perhaps you missed this thread.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2823191,
      "author_name": "Oleg Kokorin",
      "author_url": "",
      "post_date": "2024-05-19T04:39:35.117000",
      "content": "<p>I've been thinking about exactly the same questions. I'm not one of the organizers, but my guess is as follows:</p>\n<ol>\n<li>It is likely to be about 200,000 but it isn't necessarily exactly 200,000.</li>\n<li>Yes, I'm pretty sure that all neighbors are included into the test index.</li>\n<li>Since the train index does include them, we have to assume that the test index contains them as well. Which is weird given all the wording about 1975.</li>\n</ol>\n<p>Also, speaking of achieving 1.0 score: as discussed <a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981\" target=\"_blank\">here</a>, and I agree with it, the current mAP metric is being incorrectly calculated. So I won't be surprised if it gets fixed one day. The current scores will roughly be halved and 1.0 will be far away.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2823220,
          "author_name": "penguin46",
          "author_url": "",
          "post_date": "2024-05-19T04:57:44.300000",
          "content": "<p>Thanks for the reply. I generally agree with you.<br>\nRegarding the size of test_index, I am concerned that it may contain up to an additional 125,000 (2500x50) to include all the neighborhoods in test.csv, in addition to the 200,000 that result from the same generation method as train_index.</p>\n<p>Let's wait for a reply from the host.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2853226,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2024-06-03T17:07:35.250000",
              "content": "<p>The test index contains approximately 200,000 rows. </p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2853231,
              "author_name": "penguin46",
              "author_url": "",
              "post_date": "2024-06-03T17:12:38.903000",
              "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nThank you for your reply!<br>\nCould you also answer the second and third questions?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2888328,
      "author_name": "yu",
      "author_url": "",
      "post_date": "2024-06-24T18:00:25.563000",
      "content": "<p>i thought the train index has way more entries than 200,000? shouldnt it contain len(meta_data) * 50  of patents?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2858345,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-06T12:40:19.340000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2823099": "Hi @addisonhoward @sohier,\n\nThe dataset description states the following. What exactly does \"equivalent in size and setup\" mean? \n\n> train_index_patent_ids.json A list of the patents included in the Whoosh index.\n> train_index A Whoosh text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries. Only includes patents published on or after 1975.\n\nI have the following questions:\n\n1. Is the size of test_index the same as train_index, 200,000?\nThat is, does the whoosh index that is run during scoring contain the same number of patents as the train_index (exactly 200,000)?\n\n2. Are all the patents in test.csv included in test_index?\nIn other words, is it possible to achieve a score close to 1.00 if ideal queries exist? This does not seem to be achievable in train_index because no matter how train.csv is constructed, it does not contain enough neighborhoods as descrived [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/497956).\n\n3. Does test_index also include the patents prior to 1975?\nThe neighborhoods of the patent in train_index contain pre-1975 patents, which may violate the setup that \"Only includes patents published on or after 1975\". Can we assume that all test.csv neighbors are on or after 1975? Or are the test.csv neighborhoods included in test_index even if they are before 1975?",
    "2853209": "Hi @addisonhoward @sohier,\n\nCould you please answer these questions?\nSorry if you are preparing answers, but I was wondering if perhaps you missed this thread.",
    "2823191": "I've been thinking about exactly the same questions. I'm not one of the organizers, but my guess is as follows:\n\n1. It is likely to be about 200,000 but it isn't necessarily exactly 200,000.\n2. Yes, I'm pretty sure that all neighbors are included into the test index.\n3. Since the train index does include them, we have to assume that the test index contains them as well. Which is weird given all the wording about 1975.\n\nAlso, speaking of achieving 1.0 score: as discussed [here](https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/499981), and I agree with it, the current mAP metric is being incorrectly calculated. So I won't be surprised if it gets fixed one day. The current scores will roughly be halved and 1.0 will be far away.\n",
    "2888328": "i thought the train index has way more entries than 200,000? shouldnt it contain len(meta_data) * 50  of patents?",
    "2858345": ""
  }
}