{
  "id": 498241,
  "title": "A challenge without labels?",
  "url": "/competitions/uspto-explainable-ai/discussion/498241",
  "author_name": "",
  "post_date": "2024-04-27T13:12:13.983083300Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>This seems to be one of the more confusing challenges.</p>\n<p>I understand the metric and the task to write queries. However, we do not have training queries or training outputs?</p>\n<p>I have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. However, without labelled training data, this becomes an unsupervised challenge. </p>\n<p>Are we supposed to use the nearest neighbours as labels for the training indexes? If so, we still need to find queries to achieve these nearest neighbours - basically making it a huge validation set, rather than a training set?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8504824%2F8a98fdedb2ea3533af6305147abd88e0%2FScreenshot%202024-04-27%20151014.png?generation=1714223524911780&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2778952",
      "postDate": "04/27/2024 13:12:13",
      "content": "<p>This seems to be one of the more confusing challenges.</p>\n<p>I understand the metric and the task to write queries. However, we do not have training queries or training outputs?</p>\n<p>I have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. However, without labelled training data, this becomes an unsupervised challenge. </p>\n<p>Are we supposed to use the nearest neighbours as labels for the training indexes? If so, we still need to find queries to achieve these nearest neighbours - basically making it a huge validation set, rather than a training set?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8504824%2F8a98fdedb2ea3533af6305147abd88e0%2FScreenshot%202024-04-27%20151014.png?generation=1714223524911780&amp;alt=media\"></p>",
      "rawMarkdown": "This seems to be one of the more confusing challenges.\n\nI understand the metric and the task to write queries. However, we do not have training queries or training outputs?\n\nI have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. However, without labelled training data, this becomes an unsupervised challenge. \n\nAre we supposed to use the nearest neighbours as labels for the training indexes? If so, we still need to find queries to achieve these nearest neighbours - basically making it a huge validation set, rather than a training set?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8504824%2F8a98fdedb2ea3533af6305147abd88e0%2FScreenshot%202024-04-27%20151014.png?generation=1714223524911780&alt=media)",
      "votes": null
    },
    {
      "id": "2778975",
      "postDate": "04/27/2024 13:25:34",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a> <strong>even hidden 2,500 test set nearest neighbours is part of train nearest neighbours only</strong>. So, for this competition if we have Queries for the train nearest neighbours then it is mostly the partial/complete match for hidden set too.</p>\n</blockquote>\n<h1>My guess how dataset is build ( from my job knowledge as i work in patent domain )</h1>\n<blockquote>\n  <p><strong>Step 1 by Kaggle :  Select Publication Numbers</strong> =&gt; Select query for unique patent number reduce by family_id ( one publication number for each family_id - from EDA is verified ) from Google BigQuery Patent</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>Step 2 by Kaggle : Find Nearest Neighbours</strong> =&gt; For each publication find top 50 nearest neighbours embedding_v1 of size 64  using vector similarity of Google BigQuery Patent =&gt; costly so not verified with BigQuery Patent</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>Step 3 by Kaggle :  Build Queries for test set</strong> =&gt; Building search for each top results need more manual resource. So, i guess query build only happen for hidden dataset of those 2,500 nearest neighbours using boolean queries.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>for Kaggler</strong> =&gt; We need to focus first to solve training Query part then training models.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>I have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. <a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a>  - <strong>will u share the respective link if it is open sourced?</strong></p>\n</blockquote>",
      "rawMarkdown": "> @valentinwerner **even hidden 2,500 test set nearest neighbours is part of train nearest neighbours only**. So, for this competition if we have Queries for the train nearest neighbours then it is mostly the partial/complete match for hidden set too.\n\n# My guess how dataset is build ( from my job knowledge as i work in patent domain )\n\n> **Step 1 by Kaggle :  Select Publication Numbers** => Select query for unique patent number reduce by family_id ( one publication number for each family_id - from EDA is verified ) from Google BigQuery Patent\n\n---\n\n> **Step 2 by Kaggle : Find Nearest Neighbours** => For each publication find top 50 nearest neighbours embedding_v1 of size 64  using vector similarity of Google BigQuery Patent => costly so not verified with BigQuery Patent\n\n---\n\n> **Step 3 by Kaggle :  Build Queries for test set** => Building search for each top results need more manual resource. So, i guess query build only happen for hidden dataset of those 2,500 nearest neighbours using boolean queries.\n\n---\n\n> **for Kaggler** => We need to focus first to solve training Query part then training models.\n\n--- \n\n\n> I have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. @valentinwerner  - **will u share the respective link if it is open sourced?**",
      "votes": null
    },
    {
      "id": "2779022",
      "postDate": "04/27/2024 13:47:31",
      "content": "<p>Thanks for the quick answer, these means that I have not overlooked the main challenge. Also its good to know that the nearest neighbours are the actual labels. </p>\n<p>Having this as a two-step challenge makes it actually quite interesting; but the risk is that failing step 1 means that step 2 is doomed - maybe there is a way of combining both steps. I am getting quite inspired now 😉</p>\n<p>Sadly, I cannot share the code, as it is related with my master thesis which contains former PII of mine (and I don't want to leak my parents address 😀)</p>",
      "rawMarkdown": "Thanks for the quick answer, these means that I have not overlooked the main challenge. Also its good to know that the nearest neighbours are the actual labels. \n\nHaving this as a two-step challenge makes it actually quite interesting; but the risk is that failing step 1 means that step 2 is doomed - maybe there is a way of combining both steps. I am getting quite inspired now 😉\n\nSadly, I cannot share the code, as it is related with my master thesis which contains former PII of mine (and I don't want to leak my parents address 😀)",
      "votes": null
    },
    {
      "id": "2779038",
      "postDate": "04/27/2024 13:51:10",
      "content": "<p>Hahah, sure. Share any published papers related to that. Interesting to read those <a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a> </p>",
      "rawMarkdown": "Hahah, sure. Share any published papers related to that. Interesting to read those @valentinwerner",
      "votes": null
    },
    {
      "id": "2779178",
      "postDate": "04/27/2024 15:06:50",
      "content": "<p>I just double-checked: actually no PII in the doc apart from my full name 😉<br>\nThis is my master thesis - and the corresponding code that was used. The problem was different in the sense, that it is relation extraction instead of query writing. Still, the model, that is trained on writing sequences and sentences is re-trained to answer in a specific format of for example \" US  Russia  Make Public Statement\" </p>\n<p>I believe the code is over complicated, but the idea could be quite similar, once we got training queries</p>",
      "rawMarkdown": "I just double-checked: actually no PII in the doc apart from my full name 😉\nThis is my master thesis - and the corresponding code that was used. The problem was different in the sense, that it is relation extraction instead of query writing. Still, the model, that is trained on writing sequences and sentences is re-trained to answer in a specific format of for example \"<triplet> US <subj> Russia <obj> Make Public Statement\" \n\nI believe the code is over complicated, but the idea could be quite similar, once we got training queries",
      "votes": null
    },
    {
      "id": "2779184",
      "postDate": "04/27/2024 15:08:23",
      "content": "<p>Ok <a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a> Got it.</p>",
      "rawMarkdown": "Ok @valentinwerner Got it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2778975,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "04/27/2024 13:25:34",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a> <strong>even hidden 2,500 test set nearest neighbours is part of train nearest neighbours only</strong>. So, for this competition if we have Queries for the train nearest neighbours then it is mostly the partial/complete match for hidden set too.</p>\n</blockquote>\n<h1>My guess how dataset is build ( from my job knowledge as i work in patent domain )</h1>\n<blockquote>\n  <p><strong>Step 1 by Kaggle :  Select Publication Numbers</strong> =&gt; Select query for unique patent number reduce by family_id ( one publication number for each family_id - from EDA is verified ) from Google BigQuery Patent</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>Step 2 by Kaggle : Find Nearest Neighbours</strong> =&gt; For each publication find top 50 nearest neighbours embedding_v1 of size 64  using vector similarity of Google BigQuery Patent =&gt; costly so not verified with BigQuery Patent</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>Step 3 by Kaggle :  Build Queries for test set</strong> =&gt; Building search for each top results need more manual resource. So, i guess query build only happen for hidden dataset of those 2,500 nearest neighbours using boolean queries.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>for Kaggler</strong> =&gt; We need to focus first to solve training Query part then training models.</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>I have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. <a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a>  - <strong>will u share the respective link if it is open sourced?</strong></p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2779022,
          "author_name": "valentinwerner",
          "author_url": "",
          "post_date": "04/27/2024 13:47:31",
          "content": "<p>Thanks for the quick answer, these means that I have not overlooked the main challenge. Also its good to know that the nearest neighbours are the actual labels. </p>\n<p>Having this as a two-step challenge makes it actually quite interesting; but the risk is that failing step 1 means that step 2 is doomed - maybe there is a way of combining both steps. I am getting quite inspired now 😉</p>\n<p>Sadly, I cannot share the code, as it is related with my master thesis which contains former PII of mine (and I don't want to leak my parents address 😀)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2779038,
              "author_name": "seshurajup",
              "author_url": "",
              "post_date": "04/27/2024 13:51:10",
              "content": "<p>Hahah, sure. Share any published papers related to that. Interesting to read those <a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a> </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2779178,
                  "author_name": "valentinwerner",
                  "author_url": "",
                  "post_date": "04/27/2024 15:06:50",
                  "content": "<p>I just double-checked: actually no PII in the doc apart from my full name 😉<br>\nThis is my master thesis - and the corresponding code that was used. The problem was different in the sense, that it is relation extraction instead of query writing. Still, the model, that is trained on writing sequences and sentences is re-trained to answer in a specific format of for example \" US  Russia  Make Public Statement\" </p>\n<p>I believe the code is over complicated, but the idea could be quite similar, once we got training queries</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2779184,
                      "author_name": "seshurajup",
                      "author_url": "",
                      "post_date": "04/27/2024 15:08:23",
                      "content": "<p>Ok <a href=\"https://www.kaggle.com/valentinwerner\" target=\"_blank\">@valentinwerner</a> Got it.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2778952": "This seems to be one of the more confusing challenges.\n\nI understand the metric and the task to write queries. However, we do not have training queries or training outputs?\n\nI have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. However, without labelled training data, this becomes an unsupervised challenge. \n\nAre we supposed to use the nearest neighbours as labels for the training indexes? If so, we still need to find queries to achieve these nearest neighbours - basically making it a huge validation set, rather than a training set?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8504824%2F8a98fdedb2ea3533af6305147abd88e0%2FScreenshot%202024-04-27%20151014.png?generation=1714223524911780&alt=media)",
    "2778975": "> @valentinwerner **even hidden 2,500 test set nearest neighbours is part of train nearest neighbours only**. So, for this competition if we have Queries for the train nearest neighbours then it is mostly the partial/complete match for hidden set too.\n\n# My guess how dataset is build ( from my job knowledge as i work in patent domain )\n\n> **Step 1 by Kaggle :  Select Publication Numbers** => Select query for unique patent number reduce by family_id ( one publication number for each family_id - from EDA is verified ) from Google BigQuery Patent\n\n---\n\n> **Step 2 by Kaggle : Find Nearest Neighbours** => For each publication find top 50 nearest neighbours embedding_v1 of size 64  using vector similarity of Google BigQuery Patent => costly so not verified with BigQuery Patent\n\n---\n\n> **Step 3 by Kaggle :  Build Queries for test set** => Building search for each top results need more manual resource. So, i guess query build only happen for hidden dataset of those 2,500 nearest neighbours using boolean queries.\n\n---\n\n> **for Kaggler** => We need to focus first to solve training Query part then training models.\n\n--- \n\n\n> I have trained sequence-to-sequence models in the past, where a structured output like a query is trained for; this actually works decently well. @valentinwerner  - **will u share the respective link if it is open sourced?**",
    "2779022": "Thanks for the quick answer, these means that I have not overlooked the main challenge. Also its good to know that the nearest neighbours are the actual labels. \n\nHaving this as a two-step challenge makes it actually quite interesting; but the risk is that failing step 1 means that step 2 is doomed - maybe there is a way of combining both steps. I am getting quite inspired now 😉\n\nSadly, I cannot share the code, as it is related with my master thesis which contains former PII of mine (and I don't want to leak my parents address 😀)",
    "2779038": "Hahah, sure. Share any published papers related to that. Interesting to read those @valentinwerner",
    "2779178": "I just double-checked: actually no PII in the doc apart from my full name 😉\nThis is my master thesis - and the corresponding code that was used. The problem was different in the sense, that it is relation extraction instead of query writing. Still, the model, that is trained on writing sequences and sentences is re-trained to answer in a specific format of for example \"<triplet> US <subj> Russia <obj> Make Public Statement\" \n\nI believe the code is over complicated, but the idea could be quite similar, once we got training queries",
    "2779184": "Ok @valentinwerner Got it."
  },
  "source": "meta"
}