{
  "id": 541783,
  "title": "Basic Insights on Missing Values",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/541783",
  "author_name": "",
  "post_date": "2024-10-21T12:08:18.643902500Z",
  "votes": 16,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In the dataset, the first three files and part of the fourth file contain features that are completely missing across all rows. It's unlikely these are lagged or moving average features, as the missing period is too long. A more likely explanation is that these features weren't tracked initially and were only recorded starting from the middle of the fourth file. Those features are: feature_21, feature_26, feature_27, feature_31</p>\n<p>For the other files, approximately 82.5% of the rows with missing values correspond to rows where time_id is between 0 and 67 for each day. This suggests these missing values could be due to lagged or moving average features. Those features are mostly: feature_42, feature_39, feature_53, feature_50</p>\n<p>Lastly, there are a few isolated missing values that could be missing for various reasons. I've checked discussions but didn't find much information on these. If anyone discovers more, I hope they will share their insights.</p>\n<p>the code I used to find the details above: <a href=\"https://www.kaggle.com/code/aymanallawi/basic-insights-on-missing-values\" target=\"_blank\">https://www.kaggle.com/code/aymanallawi/basic-insights-on-missing-values</a></p>",
  "messages": [
    {
      "id": "3024179",
      "postDate": "10/21/2024 12:08:18",
      "content": "<p>In the dataset, the first three files and part of the fourth file contain features that are completely missing across all rows. It's unlikely these are lagged or moving average features, as the missing period is too long. A more likely explanation is that these features weren't tracked initially and were only recorded starting from the middle of the fourth file. Those features are: feature_21, feature_26, feature_27, feature_31</p>\n<p>For the other files, approximately 82.5% of the rows with missing values correspond to rows where time_id is between 0 and 67 for each day. This suggests these missing values could be due to lagged or moving average features. Those features are mostly: feature_42, feature_39, feature_53, feature_50</p>\n<p>Lastly, there are a few isolated missing values that could be missing for various reasons. I've checked discussions but didn't find much information on these. If anyone discovers more, I hope they will share their insights.</p>\n<p>the code I used to find the details above: <a href=\"https://www.kaggle.com/code/aymanallawi/basic-insights-on-missing-values\" target=\"_blank\">https://www.kaggle.com/code/aymanallawi/basic-insights-on-missing-values</a></p>",
      "rawMarkdown": "In the dataset, the first three files and part of the fourth file contain features that are completely missing across all rows. It's unlikely these are lagged or moving average features, as the missing period is too long. A more likely explanation is that these features weren't tracked initially and were only recorded starting from the middle of the fourth file. Those features are: feature_21, feature_26, feature_27, feature_31\n\nFor the other files, approximately 82.5% of the rows with missing values correspond to rows where time_id is between 0 and 67 for each day. This suggests these missing values could be due to lagged or moving average features. Those features are mostly: feature_42, feature_39, feature_53, feature_50\n\nLastly, there are a few isolated missing values that could be missing for various reasons. I've checked discussions but didn't find much information on these. If anyone discovers more, I hope they will share their insights.\n\nthe code I used to find the details above: https://www.kaggle.com/code/aymanallawi/basic-insights-on-missing-values",
      "votes": null
    },
    {
      "id": "3026893",
      "postDate": "10/24/2024 10:30:46",
      "content": "<p>The time_ids also change partway through the fourth file. (moving from 848 time_ids to 967 at date_id 698). It does look like a change of behaviour in the monitoring.</p>\n<p>Additionally, some symbols are missing at different \"date points\", as shown here:</p>\n<ul>\n<li>Symbols 4, 24 &amp; 31 are first monitored at date_id = 952</li>\n<li>Symbols 6, 18 &amp; 32 are first monitored at date_id = 1063</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F77a1c6487216dd623a267ce131082508%2Fdownload%20(1).png?generation=1729765390965821&amp;alt=media\" alt=\"\"></p>\n<p><em>From the competition description it is possible that we'll see new symbols. It might be an idea only to use time_id for ordering and feature engineering</em></p>",
      "rawMarkdown": "The time_ids also change partway through the fourth file. (moving from 848 time_ids to 967 at date_id 698). It does look like a change of behaviour in the monitoring.\n\nAdditionally, some symbols are missing at different \"date points\", as shown here:\n\n- Symbols 4, 24 & 31 are first monitored at date_id = 952\n- Symbols 6, 18 & 32 are first monitored at date_id = 1063\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F77a1c6487216dd623a267ce131082508%2Fdownload%20(1).png?generation=1729765390965821&alt=media)\n\n*From the competition description it is possible that we'll see new symbols. It might be an idea only to use time_id for ordering and feature engineering*",
      "votes": null
    },
    {
      "id": "3026988",
      "postDate": "10/24/2024 12:16:25",
      "content": "<p>Thanks for sharing. You made very good points. Using time_id and symbol_id as training features isn't safe since their values can undergo significant shifts in the test data compared to the training data specially that online learning isn't possible for this competition due to the API constraints, making data shifting very costly to the model's accuracy.</p>",
      "rawMarkdown": "Thanks for sharing. You made very good points. Using time_id and symbol_id as training features isn't safe since their values can undergo significant shifts in the test data compared to the training data specially that online learning isn't possible for this competition due to the API constraints, making data shifting very costly to the model's accuracy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3026893,
      "author_name": "paddykb",
      "author_url": "",
      "post_date": "10/24/2024 10:30:46",
      "content": "<p>The time_ids also change partway through the fourth file. (moving from 848 time_ids to 967 at date_id 698). It does look like a change of behaviour in the monitoring.</p>\n<p>Additionally, some symbols are missing at different \"date points\", as shown here:</p>\n<ul>\n<li>Symbols 4, 24 &amp; 31 are first monitored at date_id = 952</li>\n<li>Symbols 6, 18 &amp; 32 are first monitored at date_id = 1063</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F77a1c6487216dd623a267ce131082508%2Fdownload%20(1).png?generation=1729765390965821&amp;alt=media\" alt=\"\"></p>\n<p><em>From the competition description it is possible that we'll see new symbols. It might be an idea only to use time_id for ordering and feature engineering</em></p>",
      "votes": null,
      "replies": [
        {
          "id": 3026988,
          "author_name": "aymanallawi",
          "author_url": "",
          "post_date": "10/24/2024 12:16:25",
          "content": "<p>Thanks for sharing. You made very good points. Using time_id and symbol_id as training features isn't safe since their values can undergo significant shifts in the test data compared to the training data specially that online learning isn't possible for this competition due to the API constraints, making data shifting very costly to the model's accuracy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3024179": "In the dataset, the first three files and part of the fourth file contain features that are completely missing across all rows. It's unlikely these are lagged or moving average features, as the missing period is too long. A more likely explanation is that these features weren't tracked initially and were only recorded starting from the middle of the fourth file. Those features are: feature_21, feature_26, feature_27, feature_31\n\nFor the other files, approximately 82.5% of the rows with missing values correspond to rows where time_id is between 0 and 67 for each day. This suggests these missing values could be due to lagged or moving average features. Those features are mostly: feature_42, feature_39, feature_53, feature_50\n\nLastly, there are a few isolated missing values that could be missing for various reasons. I've checked discussions but didn't find much information on these. If anyone discovers more, I hope they will share their insights.\n\nthe code I used to find the details above: https://www.kaggle.com/code/aymanallawi/basic-insights-on-missing-values",
    "3026893": "The time_ids also change partway through the fourth file. (moving from 848 time_ids to 967 at date_id 698). It does look like a change of behaviour in the monitoring.\n\nAdditionally, some symbols are missing at different \"date points\", as shown here:\n\n- Symbols 4, 24 & 31 are first monitored at date_id = 952\n- Symbols 6, 18 & 32 are first monitored at date_id = 1063\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F175917%2F77a1c6487216dd623a267ce131082508%2Fdownload%20(1).png?generation=1729765390965821&alt=media)\n\n*From the competition description it is possible that we'll see new symbols. It might be an idea only to use time_id for ordering and feature engineering*",
    "3026988": "Thanks for sharing. You made very good points. Using time_id and symbol_id as training features isn't safe since their values can undergo significant shifts in the test data compared to the training data specially that online learning isn't possible for this competition due to the API constraints, making data shifting very costly to the model's accuracy."
  },
  "source": "meta"
}