{
  "id": 541163,
  "title": "How to join test df with lags df",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/541163",
  "author_name": "",
  "post_date": "2024-10-18T00:10:32.923629900Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p><code>all_data = test.join(lags, on=['date_id', 'time_id', 'symbol_id'], how='left')</code> <br>\nThis line in my predict function is giving me an error during competition submission but works fine during normal run. Any ideas on why this is happening?</p>",
  "messages": [
    {
      "id": "3020836",
      "postDate": "10/18/2024 00:10:32",
      "content": "<p><code>all_data = test.join(lags, on=['date_id', 'time_id', 'symbol_id'], how='left')</code> <br>\nThis line in my predict function is giving me an error during competition submission but works fine during normal run. Any ideas on why this is happening?</p>",
      "rawMarkdown": "`all_data = test.join(lags, on=['date_id', 'time_id', 'symbol_id'], how='left')` \nThis line in my predict function is giving me an error during competition submission but works fine during normal run. Any ideas on why this is happening?",
      "votes": null
    },
    {
      "id": "3020966",
      "postDate": "10/18/2024 05:07:37",
      "content": "<p><a href=\"https://www.kaggle.com/snehalverma10\" target=\"_blank\">@snehalverma10</a> please peruse the post <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541171\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> has explained this in detail</p>",
      "rawMarkdown": "snehalverma10 please peruse the post [here](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541171), @chumajin has explained this in detail",
      "votes": null
    },
    {
      "id": "3020971",
      "postDate": "10/18/2024 05:13:06",
      "content": "<p>During \"normal\" run we have only one data portion with time_id equals to 0. However, during the submission, lags variable may be None. And it is not None when time_id equals to 0, again. In this case, lags contains all responders data from the day before, i.e. with different time_id values.</p>\n<p>So, simple join of these two variables seems not very useful or even may lead to errors.</p>",
      "rawMarkdown": "During \"normal\" run we have only one data portion with time_id equals to 0. However, during the submission, lags variable may be None. And it is not None when time_id equals to 0, again. In this case, lags contains all responders data from the day before, i.e. with different time_id values.\n\nSo, simple join of these two variables seems not very useful or even may lead to errors.",
      "votes": null
    },
    {
      "id": "3020973",
      "postDate": "10/18/2024 05:15:13",
      "content": "<p>check whether the test set structure or column availability differs during submission. Adding checks and safeguards for missing columns or mismatches between the test and lags DataFrames should help prevent the error.</p>",
      "rawMarkdown": "check whether the test set structure or column availability differs during submission. Adding checks and safeguards for missing columns or mismatches between the test and lags DataFrames should help prevent the error.",
      "votes": null
    },
    {
      "id": "3021043",
      "postDate": "10/18/2024 06:54:56",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Thank you for referring to my notebook.</p>\n<p><a href=\"https://www.kaggle.com/snehalverma10\" target=\"_blank\">@snehalverma10</a> Maybe, you can understand from my notebook.</p>",
      "rawMarkdown": "ravi20076 Thank you for referring to my notebook.\n\n@snehalverma10 Maybe, you can understand from my notebook.",
      "votes": null
    },
    {
      "id": "3021102",
      "postDate": "10/18/2024 07:37:29",
      "content": "<p>The given test.parquet is only one batch of the data. To fully test your code, you need to use synthetic test data to simulate all corner cases that can happen in the real submission. Check my work here on how to generate synthetic data: <a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-generate-synthetic-test-data\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-generate-synthetic-test-data</a></p>",
      "rawMarkdown": "The given test.parquet is only one batch of the data. To fully test your code, you need to use synthetic test data to simulate all corner cases that can happen in the real submission. Check my work here on how to generate synthetic data: https://www.kaggle.com/code/shiyili/js24-rmf-generate-synthetic-test-data",
      "votes": null
    },
    {
      "id": "3031370",
      "postDate": "10/29/2024 16:15:58",
      "content": "<p>It works fine on normal run because it joins a single batch. During the submission, it keeps joining new batches which leads to creation of duplicate columns and they are named as column_x, column_y and goes on like that.</p>",
      "rawMarkdown": "It works fine on normal run because it joins a single batch. During the submission, it keeps joining new batches which leads to creation of duplicate columns and they are named as column_x, column_y and goes on like that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3020966,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/18/2024 05:07:37",
      "content": "<p><a href=\"https://www.kaggle.com/snehalverma10\" target=\"_blank\">@snehalverma10</a> please peruse the post <a href=\"https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541171\" target=\"_blank\">here</a>, <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> has explained this in detail</p>",
      "votes": null,
      "replies": [
        {
          "id": 3021043,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "10/18/2024 06:54:56",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Thank you for referring to my notebook.</p>\n<p><a href=\"https://www.kaggle.com/snehalverma10\" target=\"_blank\">@snehalverma10</a> Maybe, you can understand from my notebook.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3020971,
      "author_name": "kdmitrie",
      "author_url": "",
      "post_date": "10/18/2024 05:13:06",
      "content": "<p>During \"normal\" run we have only one data portion with time_id equals to 0. However, during the submission, lags variable may be None. And it is not None when time_id equals to 0, again. In this case, lags contains all responders data from the day before, i.e. with different time_id values.</p>\n<p>So, simple join of these two variables seems not very useful or even may lead to errors.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3020973,
      "author_name": "sumit08",
      "author_url": "",
      "post_date": "10/18/2024 05:15:13",
      "content": "<p>check whether the test set structure or column availability differs during submission. Adding checks and safeguards for missing columns or mismatches between the test and lags DataFrames should help prevent the error.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3021102,
      "author_name": "shiyili",
      "author_url": "",
      "post_date": "10/18/2024 07:37:29",
      "content": "<p>The given test.parquet is only one batch of the data. To fully test your code, you need to use synthetic test data to simulate all corner cases that can happen in the real submission. Check my work here on how to generate synthetic data: <a href=\"https://www.kaggle.com/code/shiyili/js24-rmf-generate-synthetic-test-data\" target=\"_blank\">https://www.kaggle.com/code/shiyili/js24-rmf-generate-synthetic-test-data</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3031370,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "10/29/2024 16:15:58",
      "content": "<p>It works fine on normal run because it joins a single batch. During the submission, it keeps joining new batches which leads to creation of duplicate columns and they are named as column_x, column_y and goes on like that.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3020836": "`all_data = test.join(lags, on=['date_id', 'time_id', 'symbol_id'], how='left')` \nThis line in my predict function is giving me an error during competition submission but works fine during normal run. Any ideas on why this is happening?",
    "3020966": "snehalverma10 please peruse the post [here](https://www.kaggle.com/competitions/jane-street-real-time-market-data-forecasting/discussion/541171), @chumajin has explained this in detail",
    "3020971": "During \"normal\" run we have only one data portion with time_id equals to 0. However, during the submission, lags variable may be None. And it is not None when time_id equals to 0, again. In this case, lags contains all responders data from the day before, i.e. with different time_id values.\n\nSo, simple join of these two variables seems not very useful or even may lead to errors.",
    "3020973": "check whether the test set structure or column availability differs during submission. Adding checks and safeguards for missing columns or mismatches between the test and lags DataFrames should help prevent the error.",
    "3021043": "ravi20076 Thank you for referring to my notebook.\n\n@snehalverma10 Maybe, you can understand from my notebook.",
    "3021102": "The given test.parquet is only one batch of the data. To fully test your code, you need to use synthetic test data to simulate all corner cases that can happen in the real submission. Check my work here on how to generate synthetic data: https://www.kaggle.com/code/shiyili/js24-rmf-generate-synthetic-test-data",
    "3031370": "It works fine on normal run because it joins a single batch. During the submission, it keeps joining new batches which leads to creation of duplicate columns and they are named as column_x, column_y and goes on like that."
  },
  "source": "meta"
}