{
  "id": 501169,
  "title": "Test set is public?",
  "url": "/competitions/uspto-explainable-ai/discussion/501169",
  "author_name": "",
  "post_date": "2024-05-08T10:26:42.348896Z",
  "votes": 21,
  "comment_count": 4,
  "views": 0,
  "content": "<blockquote>\n  <p>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n</blockquote>\n<p>Since all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> With this approach, runtime limit doesn't enforce efficient solutions.</p>",
  "messages": [
    {
      "id": "2800728",
      "postDate": "05/08/2024 10:26:42",
      "content": "<blockquote>\n  <p>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n</blockquote>\n<p>Since all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> With this approach, runtime limit doesn't enforce efficient solutions.</p>",
      "rawMarkdown": ">test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\nSince all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? @addisonhoward @sohier With this approach, runtime limit doesn't enforce efficient solutions.",
      "votes": null
    },
    {
      "id": "2800926",
      "postDate": "05/08/2024 12:16:25",
      "content": "<p>Great point! Theoretically, it should be possible to pre-compute such a lookup table, since:</p>\n<ul>\n<li>test.csv is a <strong>subset</strong> of nearest_neighbors.csv</li>\n<li>The <strong>actual metric index</strong> is hidden, but it's safe to assume (?) that the indexed patents are also a <strong>subset</strong> of all patents contained in the <code>patent_data</code> folder. </li>\n<li>While generating queries for the look up table (offline), we can make use of a huge index containing all patents. Goal would be to find queries that optimizes <strong>mAP@50</strong> against that huge index. The LB performance against the <strong>actual metric index</strong> can only be better, since the search space is smaller (assuming actual metric index is a subset of the huge index).</li>\n</ul>\n<p>However, I have no idea if this is computationally feasible.</p>",
      "rawMarkdown": "Great point! Theoretically, it should be possible to pre-compute such a lookup table, since:\n- test.csv is a **subset** of nearest_neighbors.csv\n- The **actual metric index** is hidden, but it's safe to assume (?) that the indexed patents are also a **subset** of all patents contained in the `patent_data` folder. \n- While generating queries for the look up table (offline), we can make use of a huge index containing all patents. Goal would be to find queries that optimizes **mAP@50** against that huge index. The LB performance against the **actual metric index** can only be better, since the search space is smaller (assuming actual metric index is a subset of the huge index).\n\nHowever, I have no idea if this is computationally feasible.",
      "votes": null
    },
    {
      "id": "2800935",
      "postDate": "05/08/2024 12:19:18",
      "content": "<p>There is always at least one Kaggler who spend 50K on compute while competition prize is 15K:) With some heuristics, it is feasible.</p>",
      "rawMarkdown": "There is always at least one Kaggler who spend 50K on compute while competition prize is 15K:) With some heuristics, it is feasible.",
      "votes": null
    },
    {
      "id": "2801011",
      "postDate": "05/08/2024 12:52:07",
      "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> </p>\n<blockquote>\n  <p>Since all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> With this approach, runtime limit doesn't enforce efficient solutions.</p>\n</blockquote>\n<p>-- Even it is possible to make search queries but difficult to make test environment i.e</p>\n<blockquote>\n  <p>Total number of patents in the train_index can have more than the patent numbers in the neighbours, it result more hits.<br>\n  Using only 50 tokens with 24 are operators, its more challenging to get 1 score with offline approach.</p>\n</blockquote>",
      "rawMarkdown": "aerdem4 \n> Since all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? @addisonhoward @sohier With this approach, runtime limit doesn't enforce efficient solutions.\n\n-- Even it is possible to make search queries but difficult to make test environment i.e\n> Total number of patents in the train_index can have more than the patent numbers in the neighbours, it result more hits.\n> Using only 50 tokens with 24 are operators, its more challenging to get 1 score with offline approach.",
      "votes": null
    },
    {
      "id": "2801104",
      "postDate": "05/08/2024 13:34:15",
      "content": "<p>You are welcome to precompute all queries and submit a notebook that uses a lookup; we expect some competitors will choose to do so.</p>",
      "rawMarkdown": "You are welcome to precompute all queries and submit a notebook that uses a lookup; we expect some competitors will choose to do so.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2800926,
      "author_name": "conjuring92",
      "author_url": "",
      "post_date": "05/08/2024 12:16:25",
      "content": "<p>Great point! Theoretically, it should be possible to pre-compute such a lookup table, since:</p>\n<ul>\n<li>test.csv is a <strong>subset</strong> of nearest_neighbors.csv</li>\n<li>The <strong>actual metric index</strong> is hidden, but it's safe to assume (?) that the indexed patents are also a <strong>subset</strong> of all patents contained in the <code>patent_data</code> folder. </li>\n<li>While generating queries for the look up table (offline), we can make use of a huge index containing all patents. Goal would be to find queries that optimizes <strong>mAP@50</strong> against that huge index. The LB performance against the <strong>actual metric index</strong> can only be better, since the search space is smaller (assuming actual metric index is a subset of the huge index).</li>\n</ul>\n<p>However, I have no idea if this is computationally feasible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2800935,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "05/08/2024 12:19:18",
          "content": "<p>There is always at least one Kaggler who spend 50K on compute while competition prize is 15K:) With some heuristics, it is feasible.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2801011,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "05/08/2024 12:52:07",
      "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> </p>\n<blockquote>\n  <p>Since all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> With this approach, runtime limit doesn't enforce efficient solutions.</p>\n</blockquote>\n<p>-- Even it is possible to make search queries but difficult to make test environment i.e</p>\n<blockquote>\n  <p>Total number of patents in the train_index can have more than the patent numbers in the neighbours, it result more hits.<br>\n  Using only 50 tokens with 24 are operators, its more challenging to get 1 score with offline approach.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2801104,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "05/08/2024 13:34:15",
      "content": "<p>You are welcome to precompute all queries and submit a notebook that uses a lookup; we expect some competitors will choose to do so.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2800728": ">test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\nSince all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? @addisonhoward @sohier With this approach, runtime limit doesn't enforce efficient solutions.",
    "2800926": "Great point! Theoretically, it should be possible to pre-compute such a lookup table, since:\n- test.csv is a **subset** of nearest_neighbors.csv\n- The **actual metric index** is hidden, but it's safe to assume (?) that the indexed patents are also a **subset** of all patents contained in the `patent_data` folder. \n- While generating queries for the look up table (offline), we can make use of a huge index containing all patents. Goal would be to find queries that optimizes **mAP@50** against that huge index. The LB performance against the **actual metric index** can only be better, since the search space is smaller (assuming actual metric index is a subset of the huge index).\n\nHowever, I have no idea if this is computationally feasible.",
    "2800935": "There is always at least one Kaggler who spend 50K on compute while competition prize is 15K:) With some heuristics, it is feasible.",
    "2801011": "aerdem4 \n> Since all data is public except the fact that we don't have query labels, one can precompute all queries for all possible patents and create a lookup table for the inference? I wonder if it is allowed? @addisonhoward @sohier With this approach, runtime limit doesn't enforce efficient solutions.\n\n-- Even it is possible to make search queries but difficult to make test environment i.e\n> Total number of patents in the train_index can have more than the patent numbers in the neighbours, it result more hits.\n> Using only 50 tokens with 24 are operators, its more challenging to get 1 score with offline approach.",
    "2801104": "You are welcome to precompute all queries and submit a notebook that uses a lookup; we expect some competitors will choose to do so."
  },
  "source": "meta"
}