{
  "id": 552980,
  "title": "Any hypothesis about what features 9, 10, and 11 are?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/552980",
  "author_name": "",
  "post_date": "2024-12-22T22:02:54.857490900Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>From features.csv it seems features 10, 11, and 12 are related, and they are the only integer features other than the ids.  </p>\n<p>Feature 10 looks suspiciously like some kind of month as it is bounded between 1 and 12 and if so, seems to be tied to the symbol, not the date of the trade. It's also mostly one to one except for 3 symbols, and values 8, 9, 11 never appears:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F109f90e1626689d68a4ce01538eade78%2Fdownload.png?generation=1734904052603763&amp;alt=media\" alt=\"\"></p>\n<p>Here's similar plots for 9 and 11, whose values take up a wide range but is sparse, and also mostly one to one to the symbol</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F8b9386cdbf96b79de1e0b27ac447c3b2%2Fdownload.png?generation=1734911638780899&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F6191a99a894b05753b43af7cbad673e2%2Fdownload%20(3).png?generation=1734911646865926&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3078825",
      "postDate": "12/22/2024 22:02:54",
      "content": "<p>From features.csv it seems features 10, 11, and 12 are related, and they are the only integer features other than the ids.  </p>\n<p>Feature 10 looks suspiciously like some kind of month as it is bounded between 1 and 12 and if so, seems to be tied to the symbol, not the date of the trade. It's also mostly one to one except for 3 symbols, and values 8, 9, 11 never appears:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F109f90e1626689d68a4ce01538eade78%2Fdownload.png?generation=1734904052603763&amp;alt=media\" alt=\"\"></p>\n<p>Here's similar plots for 9 and 11, whose values take up a wide range but is sparse, and also mostly one to one to the symbol</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F8b9386cdbf96b79de1e0b27ac447c3b2%2Fdownload.png?generation=1734911638780899&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F6191a99a894b05753b43af7cbad673e2%2Fdownload%20(3).png?generation=1734911646865926&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "From features.csv it seems features 10, 11, and 12 are related, and they are the only integer features other than the ids.  \n\nFeature 10 looks suspiciously like some kind of month as it is bounded between 1 and 12 and if so, seems to be tied to the symbol, not the date of the trade. It's also mostly one to one except for 3 symbols, and values 8, 9, 11 never appears:\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F109f90e1626689d68a4ce01538eade78%2Fdownload.png?generation=1734904052603763&alt=media)\n\n Here's similar plots for 9 and 11, whose values take up a wide range but is sparse, and also mostly one to one to the symbol\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F8b9386cdbf96b79de1e0b27ac447c3b2%2Fdownload.png?generation=1734911638780899&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F6191a99a894b05753b43af7cbad673e2%2Fdownload%20(3).png?generation=1734911646865926&alt=media)",
      "votes": null
    },
    {
      "id": "3078862",
      "postDate": "12/23/2024 00:00:44",
      "content": "<p>Category perhaps? Sector/Industry for that Symbol_ID for example..</p>",
      "rawMarkdown": "Category perhaps? Sector/Industry for that Symbol_ID for example..",
      "votes": null
    },
    {
      "id": "3079033",
      "postDate": "12/23/2024 06:32:29",
      "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> agreed, features 9-11 seem like peer groups and industry labels to me</p>",
      "rawMarkdown": "julianmukaj agreed, features 9-11 seem like peer groups and industry labels to me",
      "votes": null
    },
    {
      "id": "3079761",
      "postDate": "12/24/2024 05:25:25",
      "content": "<p>I try set the cardinality as the max value for category embedding, but it don't perform well in LB. So I guest the category cardinality in public test sets is not so large, or maybe only 10 or so more than the training set. But I'm not sure it's going to get much larger in private test sets.</p>",
      "rawMarkdown": "I try set the cardinality as the max value for category embedding, but it don't perform well in LB. So I guest the category cardinality in public test sets is not so large, or maybe only 10 or so more than the training set. But I'm not sure it's going to get much larger in private test sets.",
      "votes": null
    },
    {
      "id": "3079765",
      "postDate": "12/24/2024 05:29:55",
      "content": "<p>I think setting the cardinality as the max value is too high, since from the plots the categories are sparse and stable from partitions 0 to 9 with the exception of a few symbols, maybe it won't change much in the test set at all</p>",
      "rawMarkdown": "I think setting the cardinality as the max value is too high, since from the plots the categories are sparse and stable from partitions 0 to 9 with the exception of a few symbols, maybe it won't change much in the test set at all",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3078862,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "12/23/2024 00:00:44",
      "content": "<p>Category perhaps? Sector/Industry for that Symbol_ID for example..</p>",
      "votes": null,
      "replies": [
        {
          "id": 3079033,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "12/23/2024 06:32:29",
          "content": "<p><a href=\"https://www.kaggle.com/julianmukaj\" target=\"_blank\">@julianmukaj</a> agreed, features 9-11 seem like peer groups and industry labels to me</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3079761,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "12/24/2024 05:25:25",
      "content": "<p>I try set the cardinality as the max value for category embedding, but it don't perform well in LB. So I guest the category cardinality in public test sets is not so large, or maybe only 10 or so more than the training set. But I'm not sure it's going to get much larger in private test sets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3079765,
          "author_name": "redfoongus",
          "author_url": "",
          "post_date": "12/24/2024 05:29:55",
          "content": "<p>I think setting the cardinality as the max value is too high, since from the plots the categories are sparse and stable from partitions 0 to 9 with the exception of a few symbols, maybe it won't change much in the test set at all</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3078825": "From features.csv it seems features 10, 11, and 12 are related, and they are the only integer features other than the ids.  \n\nFeature 10 looks suspiciously like some kind of month as it is bounded between 1 and 12 and if so, seems to be tied to the symbol, not the date of the trade. It's also mostly one to one except for 3 symbols, and values 8, 9, 11 never appears:\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F109f90e1626689d68a4ce01538eade78%2Fdownload.png?generation=1734904052603763&alt=media)\n\n Here's similar plots for 9 and 11, whose values take up a wide range but is sparse, and also mostly one to one to the symbol\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F8b9386cdbf96b79de1e0b27ac447c3b2%2Fdownload.png?generation=1734911638780899&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640759%2F6191a99a894b05753b43af7cbad673e2%2Fdownload%20(3).png?generation=1734911646865926&alt=media)",
    "3078862": "Category perhaps? Sector/Industry for that Symbol_ID for example..",
    "3079033": "julianmukaj agreed, features 9-11 seem like peer groups and industry labels to me",
    "3079761": "I try set the cardinality as the max value for category embedding, but it don't perform well in LB. So I guest the category cardinality in public test sets is not so large, or maybe only 10 or so more than the training set. But I'm not sure it's going to get much larger in private test sets.",
    "3079765": "I think setting the cardinality as the max value is too high, since from the plots the categories are sparse and stable from partitions 0 to 9 with the exception of a few symbols, maybe it won't change much in the test set at all"
  },
  "source": "meta"
}