{
  "id": 518815,
  "title": "Please help to understand test.csv, nearest_neighbors.csv and submission.csv",
  "url": "/competitions/uspto-explainable-ai/discussion/518815",
  "author_name": "",
  "post_date": "2024-07-08T11:31:24.653743500Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Dear All,</p>\n<p>I am having difficulty understanding the .csv files mentioned in the topic title of this discussion. I have tried to understand them, and I have read all discussions to confirm my understanding, but I still have two questions.</p>\n<p>According to the competition explanation under the Data section, there is a part as follows:</p>\n<p>**test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n<p>publication_number - <strong><em><em>Only patents published on or after 1975 were included in this column.</em></em></strong><br>\ntarget_[N] - These columns specify which patents your query should yield.**</p>\n<p>A). During the submission,<br>\n(case i) Do I need to submit the queries within submission.csv file only for the samples given in test.csv (which consists of 10 samples at the moment)?<br>\nSo that, after submission, this sample test.csv (10 samples) will be replaced by 2,500 samples by the host, and generate queries using my notebook code for those 2,500 samples?</p>\n<p>OR</p>\n<p>(case ii) Do I need to randomly select 2,500 samples from nearest_neighbors.csv and generate queries and submit?</p>\n<p>B). With respect to the above question, according to sentence (Only patents published on or after 1975 were included in this column) above by USPTO, must queries be generated only for samples with publication_numbers published on or after 1975?<br>\nBecause the nearest_neighbors.csv file contains publication_numbers from before 1975 as well. <br>\nIf I randomly select 2500 samples, then ofcourse I will have samples which are before 1975. Even now from given test.csv, there are only 6 samples which are published after 1975. </p>\n<p>These two questions are really necessary for me to understand.<br>\nI sincerely request your help in understanding them.</p>\n<p>Many thanks in advance.</p>",
  "messages": [
    {
      "id": "2911549",
      "postDate": "07/08/2024 11:31:24",
      "content": "<p>Dear All,</p>\n<p>I am having difficulty understanding the .csv files mentioned in the topic title of this discussion. I have tried to understand them, and I have read all discussions to confirm my understanding, but I still have two questions.</p>\n<p>According to the competition explanation under the Data section, there is a part as follows:</p>\n<p>**test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n<p>publication_number - <strong><em><em>Only patents published on or after 1975 were included in this column.</em></em></strong><br>\ntarget_[N] - These columns specify which patents your query should yield.**</p>\n<p>A). During the submission,<br>\n(case i) Do I need to submit the queries within submission.csv file only for the samples given in test.csv (which consists of 10 samples at the moment)?<br>\nSo that, after submission, this sample test.csv (10 samples) will be replaced by 2,500 samples by the host, and generate queries using my notebook code for those 2,500 samples?</p>\n<p>OR</p>\n<p>(case ii) Do I need to randomly select 2,500 samples from nearest_neighbors.csv and generate queries and submit?</p>\n<p>B). With respect to the above question, according to sentence (Only patents published on or after 1975 were included in this column) above by USPTO, must queries be generated only for samples with publication_numbers published on or after 1975?<br>\nBecause the nearest_neighbors.csv file contains publication_numbers from before 1975 as well. <br>\nIf I randomly select 2500 samples, then ofcourse I will have samples which are before 1975. Even now from given test.csv, there are only 6 samples which are published after 1975. </p>\n<p>These two questions are really necessary for me to understand.<br>\nI sincerely request your help in understanding them.</p>\n<p>Many thanks in advance.</p>",
      "rawMarkdown": "Dear All,\n\nI am having difficulty understanding the .csv files mentioned in the topic title of this discussion. I have tried to understand them, and I have read all discussions to confirm my understanding, but I still have two questions.\n\nAccording to the competition explanation under the Data section, there is a part as follows:\n\n**test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\npublication_number - ****Only patents published on or after 1975 were included in this column.****\ntarget_[N] - These columns specify which patents your query should yield.**\n\nA). During the submission,\n(case i) Do I need to submit the queries within submission.csv file only for the samples given in test.csv (which consists of 10 samples at the moment)?\nSo that, after submission, this sample test.csv (10 samples) will be replaced by 2,500 samples by the host, and generate queries using my notebook code for those 2,500 samples?\n\nOR\n\n(case ii) Do I need to randomly select 2,500 samples from nearest_neighbors.csv and generate queries and submit?\n\nB). With respect to the above question, according to sentence (Only patents published on or after 1975 were included in this column) above by USPTO, must queries be generated only for samples with publication_numbers published on or after 1975?\nBecause the nearest_neighbors.csv file contains publication_numbers from before 1975 as well. \nIf I randomly select 2500 samples, then ofcourse I will have samples which are before 1975. Even now from given test.csv, there are only 6 samples which are published after 1975. \n\nThese two questions are really necessary for me to understand.\nI sincerely request your help in understanding them.\n\nMany thanks in advance.",
      "votes": null
    },
    {
      "id": "2911649",
      "postDate": "07/08/2024 13:03:42",
      "content": "<p>You don't have to select anything.<br>\nYour notebook is an algorithm for querying patents, but not selecting patents.</p>\n<h1>1. What is nearest_neighbors.csv for?</h1>\n<p>nearest_neighbors.csv is a csv file containing 13+ million patents and their 50 (or less) nearest neighbours. It can be used as a validation set for your query composition model (take a random 2500 patents and evaluate the quality of your model). Or even crazier, you could compose 13+ million perfect queries for each patent in nearest_neighbors.csv that found the nearest neighbours from nearest_neighbors.csv. And in the submission notebook just substitute those queries for the patents. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169</a></p>\n<h1>2. What are test.csv and submission.csv for?</h1>\n<p>The test.csv contains 10 patents and 50 (or less) neighbours to each, which ideally should be found. submission.csv - contains the same 10 patents and dummy 'text AND search' queries, which you should replace with your queries. At the time of submission, when the notebook has started to be evaluated, the number of patents is changed to 2500 unknowns and the notebook is restarted. </p>\n<h1>3. Samples after 1975</h1>\n<p>2500 unknown patents in the 'publication_number' column (test and submission).csv - will be after 1975 at submitting, but their neighbours may be found before 1975 as well</p>",
      "rawMarkdown": "You don't have to select anything.\nYour notebook is an algorithm for querying patents, but not selecting patents.\n# 1. What is nearest_neighbors.csv for?\nnearest_neighbors.csv is a csv file containing 13+ million patents and their 50 (or less) nearest neighbours. It can be used as a validation set for your query composition model (take a random 2500 patents and evaluate the quality of your model). Or even crazier, you could compose 13+ million perfect queries for each patent in nearest_neighbors.csv that found the nearest neighbours from nearest_neighbors.csv. And in the submission notebook just substitute those queries for the patents. [https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169](url)\n# 2. What are test.csv and submission.csv for?\nThe test.csv contains 10 patents and 50 (or less) neighbours to each, which ideally should be found. submission.csv - contains the same 10 patents and dummy 'text AND search' queries, which you should replace with your queries. At the time of submission, when the notebook has started to be evaluated, the number of patents is changed to 2500 unknowns and the notebook is restarted. \n# 3. Samples after 1975\n2500 unknown patents in the 'publication_number' column (test and submission).csv - will be after 1975 at submitting, but their neighbours may be found before 1975 as well",
      "votes": null
    },
    {
      "id": "2911788",
      "postDate": "07/08/2024 14:30:13",
      "content": "<p>thank you <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> for your answer. <br>\nIt is quite not clear.</p>\n<p>You said \"….will be after 1975 at filing..\" w.r.t your 3rd point. instead of \"filing date\", it should be \"publication date\" right? <br>\nbacause, that's what has been given in the details of test.csv under Data section, i.e., \"publication_number - Only patents published on or after 1975 were included in this column\".</p>\n<p>In the above context of \"patents on or after 1975\" of test.csv,  few hours ago I have made a submission, where I generated queries for only 'publication_number' which are published on or after 1975, therefore I found only 6 queries for 6 samples. Now my notebook will only generate queries for those 'publication_number' published on or after 1975 for the 2500 samples by host during evaluation.<br>\nDid I do any mistake in submission?</p>",
      "rawMarkdown": "thank you @qurusx for your answer. \nIt is quite not clear.\n\nYou said \"....will be after 1975 at filing..\" w.r.t your 3rd point. instead of \"filing date\", it should be \"publication date\" right? \nbacause, that's what has been given in the details of test.csv under Data section, i.e., \"publication_number - Only patents published on or after 1975 were included in this column\".\n\nIn the above context of \"patents on or after 1975\" of test.csv,  few hours ago I have made a submission, where I generated queries for only 'publication_number' which are published on or after 1975, therefore I found only 6 queries for 6 samples. Now my notebook will only generate queries for those 'publication_number' published on or after 1975 for the 2500 samples by host during evaluation.\nDid I do any mistake in submission?",
      "votes": null
    },
    {
      "id": "2911891",
      "postDate": "07/08/2024 15:46:51",
      "content": "<p>I meant filling-submitting notebook. I did not mean the date of filing or publication about data. <br>\nWhen you click on submit prediction in notebook editing mode, your notebook 1) runs for 10 patents and saves the new version 2) runs 2 times for 2500 patents and the notebook is evaluated. For the first point, it doesn't matter what your notebook does at all, the main thing is just running the notebook without errors and in the output submission.csv (even if it contains 6 rows). But for the 2nd point it is not only important to run your notebook without errors + but also correct submission.csv (2500)+ scoring for 1 hour. No need to filter 'publication_number' by year. The 1975+ years information is only for 2500 patents in 'publication_number'. <br>\nBut the errors during the 2nd run can be totally different and unpredictable. Lack of RAM, incorrect submission.csv, notebook running for 9+ hours, evaluation for 1+ hour and many more.<br>\nYou can skip the 1st point altogether, just submitting the notebook with output to submit directly to score<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2Fd714def2f234b7ac113999ac5ea178f0%2Fphoto_2024-07-08_18-46-22.jpg?generation=1720453584569493&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F6d5f13fc70bb6e3c59064cd8cc1c7dd3%2Fphoto_2024-07-08_18-46-29.jpg?generation=1720453591453058&amp;alt=media\"></p>",
      "rawMarkdown": "I meant filling-submitting notebook. I did not mean the date of filing or publication about data. \n\nWhen you click on submit prediction in notebook editing mode, your notebook 1) runs for 10 patents and saves the new version 2) runs 2 times for 2500 patents and the notebook is evaluated. For the first point, it doesn't matter what your notebook does at all, the main thing is just running the notebook without errors and in the output submission.csv (even if it contains 6 rows). But for the 2nd point it is not only important to run your notebook without errors + but also correct submission.csv (2500)+ scoring for 1 hour. No need to filter 'publication_number' by year. The 1975+ years information is only for 2500 patents in 'publication_number'. \n\nBut the errors during the 2nd run can be totally different and unpredictable. Lack of RAM, incorrect submission.csv, notebook running for 9+ hours, evaluation for 1+ hour and many more.\n\nYou can skip the 1st point altogether, just submitting the notebook with output to submit directly to score\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2Fd714def2f234b7ac113999ac5ea178f0%2Fphoto_2024-07-08_18-46-22.jpg?generation=1720453584569493&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F6d5f13fc70bb6e3c59064cd8cc1c7dd3%2Fphoto_2024-07-08_18-46-29.jpg?generation=1720453591453058&alt=media)",
      "votes": null
    },
    {
      "id": "2912999",
      "postDate": "07/09/2024 06:58:27",
      "content": "<p>thank you </p>",
      "rawMarkdown": "thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2911649,
      "author_name": "qurusx",
      "author_url": "",
      "post_date": "07/08/2024 13:03:42",
      "content": "<p>You don't have to select anything.<br>\nYour notebook is an algorithm for querying patents, but not selecting patents.</p>\n<h1>1. What is nearest_neighbors.csv for?</h1>\n<p>nearest_neighbors.csv is a csv file containing 13+ million patents and their 50 (or less) nearest neighbours. It can be used as a validation set for your query composition model (take a random 2500 patents and evaluate the quality of your model). Or even crazier, you could compose 13+ million perfect queries for each patent in nearest_neighbors.csv that found the nearest neighbours from nearest_neighbors.csv. And in the submission notebook just substitute those queries for the patents. <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169</a></p>\n<h1>2. What are test.csv and submission.csv for?</h1>\n<p>The test.csv contains 10 patents and 50 (or less) neighbours to each, which ideally should be found. submission.csv - contains the same 10 patents and dummy 'text AND search' queries, which you should replace with your queries. At the time of submission, when the notebook has started to be evaluated, the number of patents is changed to 2500 unknowns and the notebook is restarted. </p>\n<h1>3. Samples after 1975</h1>\n<p>2500 unknown patents in the 'publication_number' column (test and submission).csv - will be after 1975 at submitting, but their neighbours may be found before 1975 as well</p>",
      "votes": null,
      "replies": [
        {
          "id": 2911788,
          "author_name": "renukswamy",
          "author_url": "",
          "post_date": "07/08/2024 14:30:13",
          "content": "<p>thank you <a href=\"https://www.kaggle.com/qurusx\" target=\"_blank\">@qurusx</a> for your answer. <br>\nIt is quite not clear.</p>\n<p>You said \"….will be after 1975 at filing..\" w.r.t your 3rd point. instead of \"filing date\", it should be \"publication date\" right? <br>\nbacause, that's what has been given in the details of test.csv under Data section, i.e., \"publication_number - Only patents published on or after 1975 were included in this column\".</p>\n<p>In the above context of \"patents on or after 1975\" of test.csv,  few hours ago I have made a submission, where I generated queries for only 'publication_number' which are published on or after 1975, therefore I found only 6 queries for 6 samples. Now my notebook will only generate queries for those 'publication_number' published on or after 1975 for the 2500 samples by host during evaluation.<br>\nDid I do any mistake in submission?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2911891,
              "author_name": "qurusx",
              "author_url": "",
              "post_date": "07/08/2024 15:46:51",
              "content": "<p>I meant filling-submitting notebook. I did not mean the date of filing or publication about data. <br>\nWhen you click on submit prediction in notebook editing mode, your notebook 1) runs for 10 patents and saves the new version 2) runs 2 times for 2500 patents and the notebook is evaluated. For the first point, it doesn't matter what your notebook does at all, the main thing is just running the notebook without errors and in the output submission.csv (even if it contains 6 rows). But for the 2nd point it is not only important to run your notebook without errors + but also correct submission.csv (2500)+ scoring for 1 hour. No need to filter 'publication_number' by year. The 1975+ years information is only for 2500 patents in 'publication_number'. <br>\nBut the errors during the 2nd run can be totally different and unpredictable. Lack of RAM, incorrect submission.csv, notebook running for 9+ hours, evaluation for 1+ hour and many more.<br>\nYou can skip the 1st point altogether, just submitting the notebook with output to submit directly to score<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2Fd714def2f234b7ac113999ac5ea178f0%2Fphoto_2024-07-08_18-46-22.jpg?generation=1720453584569493&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F6d5f13fc70bb6e3c59064cd8cc1c7dd3%2Fphoto_2024-07-08_18-46-29.jpg?generation=1720453591453058&amp;alt=media\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2912999,
                  "author_name": "renukswamy",
                  "author_url": "",
                  "post_date": "07/09/2024 06:58:27",
                  "content": "<p>thank you </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2911549": "Dear All,\n\nI am having difficulty understanding the .csv files mentioned in the topic title of this discussion. I have tried to understand them, and I have read all discussions to confirm my understanding, but I still have two questions.\n\nAccording to the competition explanation under the Data section, there is a part as follows:\n\n**test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\npublication_number - ****Only patents published on or after 1975 were included in this column.****\ntarget_[N] - These columns specify which patents your query should yield.**\n\nA). During the submission,\n(case i) Do I need to submit the queries within submission.csv file only for the samples given in test.csv (which consists of 10 samples at the moment)?\nSo that, after submission, this sample test.csv (10 samples) will be replaced by 2,500 samples by the host, and generate queries using my notebook code for those 2,500 samples?\n\nOR\n\n(case ii) Do I need to randomly select 2,500 samples from nearest_neighbors.csv and generate queries and submit?\n\nB). With respect to the above question, according to sentence (Only patents published on or after 1975 were included in this column) above by USPTO, must queries be generated only for samples with publication_numbers published on or after 1975?\nBecause the nearest_neighbors.csv file contains publication_numbers from before 1975 as well. \nIf I randomly select 2500 samples, then ofcourse I will have samples which are before 1975. Even now from given test.csv, there are only 6 samples which are published after 1975. \n\nThese two questions are really necessary for me to understand.\nI sincerely request your help in understanding them.\n\nMany thanks in advance.",
    "2911649": "You don't have to select anything.\nYour notebook is an algorithm for querying patents, but not selecting patents.\n# 1. What is nearest_neighbors.csv for?\nnearest_neighbors.csv is a csv file containing 13+ million patents and their 50 (or less) nearest neighbours. It can be used as a validation set for your query composition model (take a random 2500 patents and evaluate the quality of your model). Or even crazier, you could compose 13+ million perfect queries for each patent in nearest_neighbors.csv that found the nearest neighbours from nearest_neighbors.csv. And in the submission notebook just substitute those queries for the patents. [https://www.kaggle.com/competitions/uspto-explainable-ai/discussion/501169](url)\n# 2. What are test.csv and submission.csv for?\nThe test.csv contains 10 patents and 50 (or less) neighbours to each, which ideally should be found. submission.csv - contains the same 10 patents and dummy 'text AND search' queries, which you should replace with your queries. At the time of submission, when the notebook has started to be evaluated, the number of patents is changed to 2500 unknowns and the notebook is restarted. \n# 3. Samples after 1975\n2500 unknown patents in the 'publication_number' column (test and submission).csv - will be after 1975 at submitting, but their neighbours may be found before 1975 as well",
    "2911788": "thank you @qurusx for your answer. \nIt is quite not clear.\n\nYou said \"....will be after 1975 at filing..\" w.r.t your 3rd point. instead of \"filing date\", it should be \"publication date\" right? \nbacause, that's what has been given in the details of test.csv under Data section, i.e., \"publication_number - Only patents published on or after 1975 were included in this column\".\n\nIn the above context of \"patents on or after 1975\" of test.csv,  few hours ago I have made a submission, where I generated queries for only 'publication_number' which are published on or after 1975, therefore I found only 6 queries for 6 samples. Now my notebook will only generate queries for those 'publication_number' published on or after 1975 for the 2500 samples by host during evaluation.\nDid I do any mistake in submission?",
    "2911891": "I meant filling-submitting notebook. I did not mean the date of filing or publication about data. \n\nWhen you click on submit prediction in notebook editing mode, your notebook 1) runs for 10 patents and saves the new version 2) runs 2 times for 2500 patents and the notebook is evaluated. For the first point, it doesn't matter what your notebook does at all, the main thing is just running the notebook without errors and in the output submission.csv (even if it contains 6 rows). But for the 2nd point it is not only important to run your notebook without errors + but also correct submission.csv (2500)+ scoring for 1 hour. No need to filter 'publication_number' by year. The 1975+ years information is only for 2500 patents in 'publication_number'. \n\nBut the errors during the 2nd run can be totally different and unpredictable. Lack of RAM, incorrect submission.csv, notebook running for 9+ hours, evaluation for 1+ hour and many more.\n\nYou can skip the 1st point altogether, just submitting the notebook with output to submit directly to score\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2Fd714def2f234b7ac113999ac5ea178f0%2Fphoto_2024-07-08_18-46-22.jpg?generation=1720453584569493&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16265464%2F6d5f13fc70bb6e3c59064cd8cc1c7dd3%2Fphoto_2024-07-08_18-46-29.jpg?generation=1720453591453058&alt=media)",
    "2912999": "thank you"
  },
  "source": "meta"
}