{
  "id": 510572,
  "title": "What are the column names in test.csv? [Solved]",
  "url": "/competitions/uspto-explainable-ai/discussion/510572",
  "author_name": "",
  "post_date": "2024-06-06T16:48:08.202290900Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>The column names in test.csv are described as below, but the following code threw an exception. It works fine for the dummy test.csv. Is the description incorrect? </p>\n<blockquote>\n  <p>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n  <p><code>publication_number</code> - Only patents published on or after 1975 were included in this column.<br>\n  <code>target_[N]</code> - These columns specify which patents your query should yield.</p>\n</blockquote>\n<p>code: (<a href=\"https://www.kaggle.com/code/ryotayoshinobu/uspto-debug/notebook\" target=\"_blank\">https://www.kaggle.com/code/ryotayoshinobu/uspto-debug/notebook</a>)</p>\n<pre><code> polars  pl\n\ntest = pl.read_csv()\ncolumns = ([] + [  i  ()])\n\n (test.columns) != (columns):\n    \n    \n (test.columns) != :\n    \n    \n:\n    pl.read_csv().write_csv()\n</code></pre>\n<p>result:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F9d21f4cb5524120503ddd3f3e92115b2%2Ferr.png?generation=1717692352834997&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2858822",
      "postDate": "06/06/2024 16:48:08",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>,</p>\n<p>The column names in test.csv are described as below, but the following code threw an exception. It works fine for the dummy test.csv. Is the description incorrect? </p>\n<blockquote>\n  <p>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</p>\n  <p><code>publication_number</code> - Only patents published on or after 1975 were included in this column.<br>\n  <code>target_[N]</code> - These columns specify which patents your query should yield.</p>\n</blockquote>\n<p>code: (<a href=\"https://www.kaggle.com/code/ryotayoshinobu/uspto-debug/notebook\" target=\"_blank\">https://www.kaggle.com/code/ryotayoshinobu/uspto-debug/notebook</a>)</p>\n<pre><code> polars  pl\n\ntest = pl.read_csv()\ncolumns = ([] + [  i  ()])\n\n (test.columns) != (columns):\n    \n    \n (test.columns) != :\n    \n    \n:\n    pl.read_csv().write_csv()\n</code></pre>\n<p>result:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F9d21f4cb5524120503ddd3f3e92115b2%2Ferr.png?generation=1717692352834997&amp;alt=media\"></p>",
      "rawMarkdown": "Hi @addisonhoward @sohier,\n\nThe column names in test.csv are described as below, but the following code threw an exception. It works fine for the dummy test.csv. Is the description incorrect? \n\n> test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\n> `publication_number` - Only patents published on or after 1975 were included in this column.\n> `target_[N]` - These columns specify which patents your query should yield.\n\ncode: (https://www.kaggle.com/code/ryotayoshinobu/uspto-debug/notebook)\n```\nimport polars as pl\n\ntest = pl.read_csv(\"/kaggle/input/uspto-explainable-ai/test.csv\")\ncolumns = set([\"publication_number\"] + [f\"target_{i}\" for i in range(50)])\n\nif set(test.columns) != set(columns):\n    # exception\n    raise\nelif len(test.columns) != 51:\n    # submission.csv not found\n    pass\nelse:\n    pl.read_csv(\"/kaggle/input/uspto-explainable-ai/sample_submission.csv\").write_csv(\"submission.csv\")\n```\n\nresult:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F9d21f4cb5524120503ddd3f3e92115b2%2Ferr.png?generation=1717692352834997&alt=media\" width=\"500\">",
      "votes": null
    },
    {
      "id": "2859782",
      "postDate": "06/07/2024 08:41:04",
      "content": "<p>The data description has several inconsistencies and at this point I've embraced that dealing with them is part of the competition.</p>\n<p>(it should not be)</p>",
      "rawMarkdown": "The data description has several inconsistencies and at this point I've embraced that dealing with them is part of the competition.\n\n(it should not be)",
      "votes": null
    },
    {
      "id": "2859812",
      "postDate": "06/07/2024 09:15:28",
      "content": "<p>I completely agree. The task is very interesting, but first we need to understand the dataset properly.</p>",
      "rawMarkdown": "I completely agree. The task is very interesting, but first we need to understand the dataset properly.",
      "votes": null
    },
    {
      "id": "2865578",
      "postDate": "06/10/2024 20:03:34",
      "content": "<p>Sorry about that. I see the issue and am uploading a patch now. It will probably take another couple of hours for the update to go live.</p>",
      "rawMarkdown": "Sorry about that. I see the issue and am uploading a patch now. It will probably take another couple of hours for the update to go live.",
      "votes": null
    },
    {
      "id": "2866158",
      "postDate": "06/11/2024 07:01:15",
      "content": "<p>Thank you! I have confirmed that the above code passes without errors.</p>",
      "rawMarkdown": "Thank you! I have confirmed that the above code passes without errors.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2859782,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "06/07/2024 08:41:04",
      "content": "<p>The data description has several inconsistencies and at this point I've embraced that dealing with them is part of the competition.</p>\n<p>(it should not be)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2859812,
          "author_name": "ryotayoshinobu",
          "author_url": "",
          "post_date": "06/07/2024 09:15:28",
          "content": "<p>I completely agree. The task is very interesting, but first we need to understand the dataset properly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2865578,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "06/10/2024 20:03:34",
      "content": "<p>Sorry about that. I see the issue and am uploading a patch now. It will probably take another couple of hours for the update to go live.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2866158,
          "author_name": "ryotayoshinobu",
          "author_url": "",
          "post_date": "06/11/2024 07:01:15",
          "content": "<p>Thank you! I have confirmed that the above code passes without errors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2858822": "Hi @addisonhoward @sohier,\n\nThe column names in test.csv are described as below, but the following code threw an exception. It works fine for the dummy test.csv. Is the description incorrect? \n\n> test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.\n\n> `publication_number` - Only patents published on or after 1975 were included in this column.\n> `target_[N]` - These columns specify which patents your query should yield.\n\ncode: (https://www.kaggle.com/code/ryotayoshinobu/uspto-debug/notebook)\n```\nimport polars as pl\n\ntest = pl.read_csv(\"/kaggle/input/uspto-explainable-ai/test.csv\")\ncolumns = set([\"publication_number\"] + [f\"target_{i}\" for i in range(50)])\n\nif set(test.columns) != set(columns):\n    # exception\n    raise\nelif len(test.columns) != 51:\n    # submission.csv not found\n    pass\nelse:\n    pl.read_csv(\"/kaggle/input/uspto-explainable-ai/sample_submission.csv\").write_csv(\"submission.csv\")\n```\n\nresult:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3962786%2F9d21f4cb5524120503ddd3f3e92115b2%2Ferr.png?generation=1717692352834997&alt=media\" width=\"500\">",
    "2859782": "The data description has several inconsistencies and at this point I've embraced that dealing with them is part of the competition.\n\n(it should not be)",
    "2859812": "I completely agree. The task is very interesting, but first we need to understand the dataset properly.",
    "2865578": "Sorry about that. I see the issue and am uploading a patch now. It will probably take another couple of hours for the update to go live.",
    "2866158": "Thank you! I have confirmed that the above code passes without errors."
  },
  "source": "meta"
}