{
  "id": 546160,
  "title": "Do you even do data engineering?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/546160",
  "author_name": "",
  "post_date": "2024-11-14T07:51:14.665923Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Its a huge dataset. feature engineering would make it even larger.<br>\nI got a 128g RAM, but still having a hard time managing the data.<br>\nDo people even do feature engineering? I tried lags but the results become even worse…..</p>",
  "messages": [
    {
      "id": "3045139",
      "postDate": "11/14/2024 07:51:14",
      "content": "<p>Its a huge dataset. feature engineering would make it even larger.<br>\nI got a 128g RAM, but still having a hard time managing the data.<br>\nDo people even do feature engineering? I tried lags but the results become even worse…..</p>",
      "rawMarkdown": "Its a huge dataset. feature engineering would make it even larger.\nI got a 128g RAM, but still having a hard time managing the data.\nDo people even do feature engineering? I tried lags but the results become even worse.....",
      "votes": null
    },
    {
      "id": "3045141",
      "postDate": "11/14/2024 07:54:11",
      "content": "<p><a href=\"https://www.kaggle.com/zoutain\" target=\"_blank\">@zoutain</a> this is a very difficult competition due to the below-</p>\n<ol>\n<li>Choice of dates and training period is highly influential to the CV and Lb score</li>\n<li>It is very easy to develop a model, but quite hard to develop a <strong>robust model</strong></li>\n<li>Conventional approaches are not likely to yield any satisfactory result</li>\n<li>Normal CV approaches are most unlikely to correlate with the LB</li>\n<li>Data size, submission 1 minute problem and training constraints are making it even more difficult to build a good process</li>\n<li>High scoring public approaches are extremely unlikely to hold water beyond the public LB period. Score clusters created this way are highly misleading </li>\n</ol>\n<p>I have a 128GB RAM PC and it is more than enough for me at this stage. </p>",
      "rawMarkdown": "zoutain this is a very difficult competition due to the below-\n1. Choice of dates and training period is highly influential to the CV and Lb score\n2. It is very easy to develop a model, but quite hard to develop a **robust model**\n3. Conventional approaches are not likely to yield any satisfactory result\n4. Normal CV approaches are most unlikely to correlate with the LB\n5. Data size, submission 1 minute problem and training constraints are making it even more difficult to build a good process\n6. High scoring public approaches are extremely unlikely to hold water beyond the public LB period. Score clusters created this way are highly misleading \n\nI have a 128GB RAM PC and it is more than enough for me at this stage.",
      "votes": null
    },
    {
      "id": "3045145",
      "postDate": "11/14/2024 07:59:46",
      "content": "<p>good to know.</p>",
      "rawMarkdown": "good to know.",
      "votes": null
    },
    {
      "id": "3045598",
      "postDate": "11/14/2024 16:05:42",
      "content": "<p>Yes. If I had 128gb of RAM I'd be BALLIN right now</p>",
      "rawMarkdown": "Yes. If I had 128gb of RAM I'd be BALLIN right now",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3045141,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "11/14/2024 07:54:11",
      "content": "<p><a href=\"https://www.kaggle.com/zoutain\" target=\"_blank\">@zoutain</a> this is a very difficult competition due to the below-</p>\n<ol>\n<li>Choice of dates and training period is highly influential to the CV and Lb score</li>\n<li>It is very easy to develop a model, but quite hard to develop a <strong>robust model</strong></li>\n<li>Conventional approaches are not likely to yield any satisfactory result</li>\n<li>Normal CV approaches are most unlikely to correlate with the LB</li>\n<li>Data size, submission 1 minute problem and training constraints are making it even more difficult to build a good process</li>\n<li>High scoring public approaches are extremely unlikely to hold water beyond the public LB period. Score clusters created this way are highly misleading </li>\n</ol>\n<p>I have a 128GB RAM PC and it is more than enough for me at this stage. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3045145,
          "author_name": "zoutain",
          "author_url": "",
          "post_date": "11/14/2024 07:59:46",
          "content": "<p>good to know.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3045598,
      "author_name": "jackvd",
      "author_url": "",
      "post_date": "11/14/2024 16:05:42",
      "content": "<p>Yes. If I had 128gb of RAM I'd be BALLIN right now</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3045139": "Its a huge dataset. feature engineering would make it even larger.\nI got a 128g RAM, but still having a hard time managing the data.\nDo people even do feature engineering? I tried lags but the results become even worse.....",
    "3045141": "zoutain this is a very difficult competition due to the below-\n1. Choice of dates and training period is highly influential to the CV and Lb score\n2. It is very easy to develop a model, but quite hard to develop a **robust model**\n3. Conventional approaches are not likely to yield any satisfactory result\n4. Normal CV approaches are most unlikely to correlate with the LB\n5. Data size, submission 1 minute problem and training constraints are making it even more difficult to build a good process\n6. High scoring public approaches are extremely unlikely to hold water beyond the public LB period. Score clusters created this way are highly misleading \n\nI have a 128GB RAM PC and it is more than enough for me at this stage.",
    "3045145": "good to know.",
    "3045598": "Yes. If I had 128gb of RAM I'd be BALLIN right now"
  },
  "source": "meta"
}