{
  "id": 543212,
  "title": "What is test set's 'row_id'? ",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/543212",
  "author_name": "T M",
  "post_date": "2024-10-29T10:21:17.391000",
  "votes": 1,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I'm probably missing something here: what is the 'row_id' found in the test set and not in the training set? How is it obtained, based on training set data?</p>",
  "messages": [
    {
      "id": 3031980,
      "postDate": "2024-10-30T11:45:45.163Z",
      "content": "<p>row_id is symbol_id?</p>",
      "rawMarkdown": "row_id is symbol_id?",
      "votes": 1,
      "replies": [
        {
          "id": 3032002,
          "postDate": "2024-10-30T12:43:24.290Z",
          "content": "<p>That's the point…</p>",
          "rawMarkdown": "That's the point..."
        }
      ]
    },
    {
      "id": 3031093,
      "postDate": "2024-10-29T10:21:17.390Z",
      "content": "<p>I'm probably missing something here: what is the 'row_id' found in the test set and not in the training set? How is it obtained, based on training set data?</p>",
      "rawMarkdown": "I'm probably missing something here: what is the 'row_id' found in the test set and not in the training set? How is it obtained, based on training set data?",
      "votes": 1
    },
    {
      "id": 3031106,
      "postDate": "2024-10-29T10:45:04.913Z",
      "content": "<p>It has nothing to do with the training data. It’s simply a row index to make sure your predictions align correctly with the test batch.</p>",
      "rawMarkdown": "It has nothing to do with the training data. It’s simply a row index to make sure your predictions align correctly with the test batch.",
      "replies": [
        {
          "id": 3031152,
          "postDate": "2024-10-29T12:23:23.887Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>. But it does matter A LOT, as every 'symbol_id' is an individual asset (\"identifies a unique financial instrument\"). </p>",
          "rawMarkdown": "Thanks @shiyili. But it does matter A LOT, as every 'symbol_id' is an individual asset (\"identifies a unique financial instrument\"). ",
          "votes": -1,
          "replies": [
            {
              "id": 3031175,
              "postDate": "2024-10-29T12:54:30.020Z",
              "content": "<p>Isn't the question about <code>row_id</code> ?</p>",
              "rawMarkdown": "Isn't the question about `row_id` ?"
            },
            {
              "id": 3031765,
              "postDate": "2024-10-30T04:44:44.873Z",
              "content": "<p>It is. Actually, 'row_id' has to have a meaning. How do you identify an individual asset in the test set? Not by a 'row_id' absent in the training. Each 'symbol_id' contains 1 time series (with dates and times identifiers). The test set should contain the same information for each asset. Unless I'm missing something. </p>",
              "rawMarkdown": "It is. Actually, 'row_id' has to have a meaning. How do you identify an individual asset in the test set? Not by a 'row_id' absent in the training. Each 'symbol_id' contains 1 time series (with dates and times identifiers). The test set should contain the same information for each asset. Unless I'm missing something. "
            },
            {
              "id": 3031982,
              "postDate": "2024-10-30T11:46:20.977Z",
              "content": "<blockquote>\n  <p>How do you identify an individual asset in the test set?</p>\n</blockquote>\n<p>By symbol_id</p>",
              "rawMarkdown": ">How do you identify an individual asset in the test set?\n\nBy symbol_id"
            },
            {
              "id": 3032003,
              "postDate": "2024-10-30T12:43:59.210Z",
              "content": "<p>symbol_id = row_id?</p>",
              "rawMarkdown": "symbol_id = row_id?"
            },
            {
              "id": 3032026,
              "postDate": "2024-10-30T13:14:46.817Z",
              "content": "<p>No <a href=\"https://www.kaggle.com/thierryme\" target=\"_blank\">@thierryme</a> . simply print a test.parquet and you won't be so confused. <br>\nsymbol_id is symbol_id, data_id is date_id, time_id is time-id. row_id is row index.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F336770a8fd5f4d451e7fa9aad91dbc0e%2FScreenshot%202024-10-30%20141225.jpg?generation=1730293996962899&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "No @thierryme . simply print a test.parquet and you won't be so confused. \nsymbol_id is symbol_id, data_id is date_id, time_id is time-id. row_id is row index.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F336770a8fd5f4d451e7fa9aad91dbc0e%2FScreenshot%202024-10-30%20141225.jpg?generation=1730293996962899&alt=media)"
            },
            {
              "id": 3032029,
              "postDate": "2024-10-30T13:21:11.243Z",
              "content": "<p>No. Look at <code>test.parquet</code>. <code>row_id</code> and <code>symbol_id</code> are clearly different columns. They are only equal because the data is sorted by <code>date_id</code>, <code>time_id</code>, and <code>symbol_id</code>. For a hypothetical <code>row_id</code>=40 (if there are no new <code>symbol_id</code>), <code>symbol_id</code> would theoretically be 0 again.</p>",
              "rawMarkdown": "No. Look at `test.parquet`. `row_id` and `symbol_id` are clearly different columns. They are only equal because the data is sorted by `date_id`, `time_id`, and `symbol_id`. For a hypothetical `row_id`=40 (if there are no new `symbol_id`), `symbol_id` would theoretically be 0 again.",
              "votes": 1
            },
            {
              "id": 3032119,
              "postDate": "2024-10-30T14:59:50.713Z",
              "content": "<p>Can you break this down? Do not think in a supervised tabular learning way, and predict everything as a bulk. Where is the asset identifier? How do you identify the future dates for EACH asset? </p>\n<pre><code>predictions = test.select(\n        ,\n        pl.lit().alias(),\n    )\n\n    \n     (predictions, pl.DataFrame | pd.DataFrame)\n    \n     predictions.columns == [, ]\n    \n     (predictions) == (test)\n</code></pre>",
              "rawMarkdown": "Can you break this down? Do not think in a supervised tabular learning way, and predict everything as a bulk. Where is the asset identifier? How do you identify the future dates for EACH asset? \n\n```python\npredictions = test.select(\n        'row_id',\n        pl.lit(0.0).alias('responder_6'),\n    )\n\n    # The predict function must return a DataFrame\n    assert isinstance(predictions, pl.DataFrame | pd.DataFrame)\n    # with columns 'row_id', 'responer_6'\n    assert predictions.columns == ['row_id', 'responder_6']\n    # and as many rows as the test data.\n    assert len(predictions) == len(test)\n```"
            },
            {
              "id": 3032878,
              "postDate": "2024-10-31T13:56:31.997Z",
              "content": "<p>That's exactly what <code>row_id</code> is for. It's to identify which prediction corresponds to which row of data. The <code>symbol_id</code> is for you to identify which symbol it is. </p>\n<p>The idea is that you take the test data, format it into the required formats, make your predictions, and then the <code>row_id</code> is used to identify which prediction corresponds to which observation.</p>\n<p>Think about it this way, instead of the function passing in <code>date_id</code>,<code>time_id</code>,<code>symbol_id</code> and <code>responder_6</code>, the function only has to send in <code>row_id</code> and <code>responder_6</code> to the server.</p>",
              "rawMarkdown": "That's exactly what `row_id` is for. It's to identify which prediction corresponds to which row of data. The `symbol_id` is for you to identify which symbol it is. \n\nThe idea is that you take the test data, format it into the required formats, make your predictions, and then the `row_id` is used to identify which prediction corresponds to which observation.\n\nThink about it this way, instead of the function passing in `date_id`,`time_id`,`symbol_id` and `responder_6`, the function only has to send in `row_id` and `responder_6` to the server."
            },
            {
              "id": 3034471,
              "postDate": "2024-11-02T07:10:10.247Z",
              "content": "<p><a href=\"https://www.kaggle.com/abstrphil\" target=\"_blank\">@abstrphil</a> Thanks, so we assume everything is always sorted </p>",
              "rawMarkdown": "@abstrphil Thanks, so we assume everything is always sorted "
            },
            {
              "id": 3034663,
              "postDate": "2024-11-02T12:52:51.813Z",
              "content": "<p>You don't have to assume anything, that's the point of <code>row_id</code>. Any two rows with the same <code>row_id</code> correspond to the same observation.</p>",
              "rawMarkdown": "You don't have to assume anything, that's the point of `row_id`. Any two rows with the same `row_id` correspond to the same observation."
            }
          ]
        }
      ]
    },
    {
      "id": 3032009,
      "postDate": "2024-10-30T12:55:30.573Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Hi, could you please address this? Need to link symbol_id, date_id and time_id in training set to symbol_id, date_id and time_id (historical data for each financial asset) in test set (future data for each financial asset)</p>",
      "rawMarkdown": "@sohier Hi, could you please address this? Need to link symbol_id, date_id and time_id in training set to symbol_id, date_id and time_id (historical data for each financial asset) in test set (future data for each financial asset)",
      "votes": -1
    },
    {
      "id": 3031139,
      "postDate": "2024-10-29T11:59:48.757Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3031153,
          "postDate": "2024-10-29T12:26:25.287Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/anthonytherrien\" target=\"_blank\">@anthonytherrien</a>! Any idea on how to get 'date_id', 'symbol_id' and 'time_id' (it does matter) from the public leaderboard? It \"usually\" comes in a file called submission_example.</p>\n<p>\"to help link predictions back to the original rows\": this might be what I'm looking for. Could you please elaborate in this context?</p>",
          "rawMarkdown": "Thanks @anthonytherrien! Any idea on how to get 'date_id', 'symbol_id' and 'time_id' (it does matter) from the public leaderboard? It \"usually\" comes in a file called submission_example.\n\n\"to help link predictions back to the original rows\": this might be what I'm looking for. Could you please elaborate in this context?",
          "replies": [
            {
              "id": 3031766,
              "postDate": "2024-10-30T04:45:08.007Z",
              "content": "<p>Actually, 'row_id' has to have a meaning. How do you identify an individual asset in the test set? Not by a 'row_id' absent in the training. Each 'symbol_id' contains 1 time series (with dates and times identifiers). The test set should contain the same information for each asset. Unless I'm missing something. </p>",
              "rawMarkdown": "Actually, 'row_id' has to have a meaning. How do you identify an individual asset in the test set? Not by a 'row_id' absent in the training. Each 'symbol_id' contains 1 time series (with dates and times identifiers). The test set should contain the same information for each asset. Unless I'm missing something. "
            },
            {
              "id": 3075676,
              "postDate": "2024-12-19T06:05:25.027Z",
              "content": "<p>Just to be very sure, just like in the sample submission file, does it start from 0 and go all the to N number of predictions?</p>",
              "rawMarkdown": "Just to be very sure, just like in the sample submission file, does it start from 0 and go all the to N number of predictions?"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3031980,
      "author_name": "JixuanZhang",
      "author_url": "",
      "post_date": "2024-10-30T11:45:45.163000",
      "content": "<p>row_id is symbol_id?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3032002,
          "author_name": "T M",
          "author_url": "",
          "post_date": "2024-10-30T12:43:24.290000",
          "content": "<p>That's the point…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3031106,
      "author_name": "SLi",
      "author_url": "",
      "post_date": "2024-10-29T10:45:04.913000",
      "content": "<p>It has nothing to do with the training data. It’s simply a row index to make sure your predictions align correctly with the test batch.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3031152,
          "author_name": "T M",
          "author_url": "",
          "post_date": "2024-10-29T12:23:23.887000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/shiyili\" target=\"_blank\">@shiyili</a>. But it does matter A LOT, as every 'symbol_id' is an individual asset (\"identifies a unique financial instrument\"). </p>",
          "votes": -1,
          "replies": [
            {
              "id": 3031175,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-10-29T12:54:30.020000",
              "content": "<p>Isn't the question about <code>row_id</code> ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3031765,
              "author_name": "T M",
              "author_url": "",
              "post_date": "2024-10-30T04:44:44.873000",
              "content": "<p>It is. Actually, 'row_id' has to have a meaning. How do you identify an individual asset in the test set? Not by a 'row_id' absent in the training. Each 'symbol_id' contains 1 time series (with dates and times identifiers). The test set should contain the same information for each asset. Unless I'm missing something. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3031982,
              "author_name": "Evgeniia Grigoreva",
              "author_url": "",
              "post_date": "2024-10-30T11:46:20.977000",
              "content": "<blockquote>\n  <p>How do you identify an individual asset in the test set?</p>\n</blockquote>\n<p>By symbol_id</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3032003,
              "author_name": "T M",
              "author_url": "",
              "post_date": "2024-10-30T12:43:59.210000",
              "content": "<p>symbol_id = row_id?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3032026,
              "author_name": "SLi",
              "author_url": "",
              "post_date": "2024-10-30T13:14:46.817000",
              "content": "<p>No <a href=\"https://www.kaggle.com/thierryme\" target=\"_blank\">@thierryme</a> . simply print a test.parquet and you won't be so confused. <br>\nsymbol_id is symbol_id, data_id is date_id, time_id is time-id. row_id is row index.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F962168%2F336770a8fd5f4d451e7fa9aad91dbc0e%2FScreenshot%202024-10-30%20141225.jpg?generation=1730293996962899&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3032029,
              "author_name": "Abstr Phil",
              "author_url": "",
              "post_date": "2024-10-30T13:21:11.243000",
              "content": "<p>No. Look at <code>test.parquet</code>. <code>row_id</code> and <code>symbol_id</code> are clearly different columns. They are only equal because the data is sorted by <code>date_id</code>, <code>time_id</code>, and <code>symbol_id</code>. For a hypothetical <code>row_id</code>=40 (if there are no new <code>symbol_id</code>), <code>symbol_id</code> would theoretically be 0 again.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3032119,
              "author_name": "T M",
              "author_url": "",
              "post_date": "2024-10-30T14:59:50.713000",
              "content": "<p>Can you break this down? Do not think in a supervised tabular learning way, and predict everything as a bulk. Where is the asset identifier? How do you identify the future dates for EACH asset? </p>\n<pre><code>predictions = test.select(\n        ,\n        pl.lit().alias(),\n    )\n\n    \n     (predictions, pl.DataFrame | pd.DataFrame)\n    \n     predictions.columns == [, ]\n    \n     (predictions) == (test)\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3032878,
              "author_name": "Abstr Phil",
              "author_url": "",
              "post_date": "2024-10-31T13:56:31.997000",
              "content": "<p>That's exactly what <code>row_id</code> is for. It's to identify which prediction corresponds to which row of data. The <code>symbol_id</code> is for you to identify which symbol it is. </p>\n<p>The idea is that you take the test data, format it into the required formats, make your predictions, and then the <code>row_id</code> is used to identify which prediction corresponds to which observation.</p>\n<p>Think about it this way, instead of the function passing in <code>date_id</code>,<code>time_id</code>,<code>symbol_id</code> and <code>responder_6</code>, the function only has to send in <code>row_id</code> and <code>responder_6</code> to the server.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3034471,
              "author_name": "T M",
              "author_url": "",
              "post_date": "2024-11-02T07:10:10.247000",
              "content": "<p><a href=\"https://www.kaggle.com/abstrphil\" target=\"_blank\">@abstrphil</a> Thanks, so we assume everything is always sorted </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3034663,
              "author_name": "Abstr Phil",
              "author_url": "",
              "post_date": "2024-11-02T12:52:51.813000",
              "content": "<p>You don't have to assume anything, that's the point of <code>row_id</code>. Any two rows with the same <code>row_id</code> correspond to the same observation.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3032009,
      "author_name": "T M",
      "author_url": "",
      "post_date": "2024-10-30T12:55:30.573000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Hi, could you please address this? Need to link symbol_id, date_id and time_id in training set to symbol_id, date_id and time_id (historical data for each financial asset) in test set (future data for each financial asset)</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3031139,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-29T11:59:48.757000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3031153,
          "author_name": "T M",
          "author_url": "",
          "post_date": "2024-10-29T12:26:25.287000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/anthonytherrien\" target=\"_blank\">@anthonytherrien</a>! Any idea on how to get 'date_id', 'symbol_id' and 'time_id' (it does matter) from the public leaderboard? It \"usually\" comes in a file called submission_example.</p>\n<p>\"to help link predictions back to the original rows\": this might be what I'm looking for. Could you please elaborate in this context?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3031766,
              "author_name": "T M",
              "author_url": "",
              "post_date": "2024-10-30T04:45:08.007000",
              "content": "<p>Actually, 'row_id' has to have a meaning. How do you identify an individual asset in the test set? Not by a 'row_id' absent in the training. Each 'symbol_id' contains 1 time series (with dates and times identifiers). The test set should contain the same information for each asset. Unless I'm missing something. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3075676,
              "author_name": "kaso",
              "author_url": "",
              "post_date": "2024-12-19T06:05:25.027000",
              "content": "<p>Just to be very sure, just like in the sample submission file, does it start from 0 and go all the to N number of predictions?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3031980": "row_id is symbol_id?",
    "3031093": "I'm probably missing something here: what is the 'row_id' found in the test set and not in the training set? How is it obtained, based on training set data?",
    "3031106": "It has nothing to do with the training data. It’s simply a row index to make sure your predictions align correctly with the test batch.",
    "3032009": "@sohier Hi, could you please address this? Need to link symbol_id, date_id and time_id in training set to symbol_id, date_id and time_id (historical data for each financial asset) in test set (future data for each financial asset)",
    "3031139": ""
  }
}