{
  "id": 596478,
  "title": "A Short Crypto Alpha Chasing Race (40th)",
  "url": "/competitions/drw-crypto-market-prediction/discussion/596478",
  "author_name": "durvorezbariq",
  "post_date": "2025-08-03T20:44:26.781000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This is a story of curiosity, learning and experimenting with one main goal — chasing alpha through the fog of crypto noise.</p>\n<p><strong>Chapter 1. Quantifying the (enemy) noise</strong><br>\nAs we all know, crypto data is notoriously known for having large noise to signal ratio. So the first thing I wanted to check is what exactly noise is in our case. A useful trick I've found is just to shoot the imaginary enemy (crypto label) at random. For example producing 100 random forecast each with of pure random numbers (<code>np.random.rand</code>, <code>np.random.norm</code>) gave a Pearson's correlations anywhere between  -0.005 to 0.006 - random walk, as expected. This essentially translates to: if your CV score moves at the 3rd decimal, you’re just spinning your wheels\". Only 2nd decimal point jumps are actual progress. </p>\n<p><strong>Chapter 2. Setup a solid CV</strong><br>\nI decided to use four folds time series CV setup.  Each time I trained on historical data and predicted the remaining future months.  I used the following starting forecasting date for the four folds respectively <code>FCT_DATES = [\"2023-08-31\", \"2023-09-30\", \"2023-10-31\", \"2023-11-30\"]</code> Essentially the first fold is trained from \"2023-03-01 to 2023-08-31\" and forecasts the remaining six months from \"2023-08-31 to 2024-02-29\". You get the idea of how the remaining train/test folds work. The key (and most time consumed):  having a modular pipeline allowing for quick iteration — data engineering, feature engineering, and modeling.</p>\n<p><strong>Chapter 3. Getting into the race</strong><br>\nI've observed there is poor correlation between the public LB and my CV.  For example making a forecast that just takes the value of the most correlated feature (X752) produced 0.09 correlation in CV and only 0.06 in the public LB. Often when I add features in CV and correlation increase, the LB correlation would actually decrease. This meant that:</p>\n<ul>\n<li>LB data is extremely noisy and results should be taken with a grain of salt</li>\n<li>Crypto data is very regime-dependent and the public LB score captures a different regime compared to the four folds in my local CV score. For example, one could clearly see how the label distribution is different in the first half of 2023 and  the second half 2023</li>\n</ul>\n<p>Having these two observations, they translated to these actions:</p>\n<ul>\n<li>Make a robust forecast via bagging techniques</li>\n<li>Since we are allowed to make two shots (aka submit two notebooks), then shoot one that aims the LB and one that aims the CV</li>\n</ul>\n<p><strong>Chapter 4. Racing</strong><br>\nI decided to use the classic AK-47, i.e a simple linear regression with forward feature selection. Most of the features had very low correlation with the label so it did not make sense to start with all of them and remove one by one via backward selection. My key racing tips and discoveries are:</p>\n<ul>\n<li>Whenever I added new features (whether one at a time or in groups), I only kept them if they improved correlation by more than the third decimal—otherwise, I considered it just noise and didn’t include them.</li>\n<li>Violating linear regression assumptions isn’t necessarily a problem for forecasting. In practice, accuracy matters more than meeting textbook criteria like p-values. For example, variables with large p-values are often highly correlated with other variables (multicollinearity), making their individual contributions hard to interpret. But including such variables doesn’t always hurt accuracy. If you compare a model with just x1 (y ~ x1) to one with both x1 and a highly correlated x2 (y ~ x1 + x2 where x2 is essentially a linear function of x1), both models can forecast equally well—sometimes the multicollinear model is even better. This happens because many different combinations of coefficients can fit the data similarly well; the model just splits “credit” between the variables.  </li>\n<li>Because of this, I optimized for the highest CV score, not spending too much time on p-values or fixing multicollinearity. The only caveat is that too many multicollinear features can make parameters unstable, and if test-time relationships change, forecasts can suffer. Bagging (averaging) helps reduce this instability.</li>\n<li>Smoothing and removing high-leverage points had the biggest impact on model beta stability and improved both leaderboard and CV scores. In crypto, high-leverage outliers pull the regression line much more than in traditional data, so handling them is crucial.  </li>\n<li>One key feature that improved my score was adding the leverage values like this:</li>\n</ul>\n<pre><code>XtX_inv = np.linalg.pinv(X_design.T @ X_design)\nM = X_design @ XtX_inv\nleverages = np.sum(M * X_design, axis=1)\ndata = data.with_columns(\n        pl.from_numpy(leverages, schema=[]).to_series()\n  )\n</code></pre>\n<p>This feature is the 'leverage' feature and essentially models the fat tails, which in crypto happen often.</p>\n<ul>\n<li>My final linear regression model used 22 features: the 15 most correlated with the label, one leverage feature, and six imbalance features.</li>\n<li>Clipping or winsorizing extreme label values did not improve CV score—in fact, it often made it worse. For example, in one of my experiments, I removed the top 2% most extreme labels and did not include them in evaluation (i.e. the test set). However, I left the extreme values in the train set. The Pearson's correlation on test set improved by 0.015. This suggests that those extreme y-values contain real signal and help guide the model in the right direction; their associated features are actually useful. That is if I could predict accurately the extreme values there was about 0.015 score boost to gain!</li>\n</ul>\n<p><strong>Chapter 5. Late laps, and racing ethics</strong><br>\nThroughout the whole journey I focused exclusively on data exploration, feature engineering and modelling. I refrained myself of peeking or attempting to reverse engineer the problem. Towards the finish, I doubled down on feature engineering — the real spirit of the competition (and makes the race much more fun and exciting). Power features, logs, imbalance measures, clipping, all tried; only the leverage feature and regime-aware engineering stuck. Clipping the most extreme labels actually hurt true test performance: those wild points, I realized, were rich with signal if I wanted to win. </p>\n<p>Final lessons for fellow racers:   </p>\n<ul>\n<li>Simple, robust models, tested on solid, time-series CV can outlast overfit magic tricks on tabular, well structured data with lots of noise.</li>\n<li>Noisy data means moving faster isn’t enough—you need to change direction fast when the regime does.</li>\n</ul>\n<p>The checkered flag may have gone down and the race over, but I already have a keen sense for the next time the starting lights go green (hopefully with a time-series API <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a>)</p>",
  "messages": [
    {
      "id": 3262553,
      "postDate": "2025-08-03T20:44:26.783Z",
      "content": "<p>This is a story of curiosity, learning and experimenting with one main goal — chasing alpha through the fog of crypto noise.</p>\n<p><strong>Chapter 1. Quantifying the (enemy) noise</strong><br>\nAs we all know, crypto data is notoriously known for having large noise to signal ratio. So the first thing I wanted to check is what exactly noise is in our case. A useful trick I've found is just to shoot the imaginary enemy (crypto label) at random. For example producing 100 random forecast each with of pure random numbers (<code>np.random.rand</code>, <code>np.random.norm</code>) gave a Pearson's correlations anywhere between  -0.005 to 0.006 - random walk, as expected. This essentially translates to: if your CV score moves at the 3rd decimal, you’re just spinning your wheels\". Only 2nd decimal point jumps are actual progress. </p>\n<p><strong>Chapter 2. Setup a solid CV</strong><br>\nI decided to use four folds time series CV setup.  Each time I trained on historical data and predicted the remaining future months.  I used the following starting forecasting date for the four folds respectively <code>FCT_DATES = [\"2023-08-31\", \"2023-09-30\", \"2023-10-31\", \"2023-11-30\"]</code> Essentially the first fold is trained from \"2023-03-01 to 2023-08-31\" and forecasts the remaining six months from \"2023-08-31 to 2024-02-29\". You get the idea of how the remaining train/test folds work. The key (and most time consumed):  having a modular pipeline allowing for quick iteration — data engineering, feature engineering, and modeling.</p>\n<p><strong>Chapter 3. Getting into the race</strong><br>\nI've observed there is poor correlation between the public LB and my CV.  For example making a forecast that just takes the value of the most correlated feature (X752) produced 0.09 correlation in CV and only 0.06 in the public LB. Often when I add features in CV and correlation increase, the LB correlation would actually decrease. This meant that:</p>\n<ul>\n<li>LB data is extremely noisy and results should be taken with a grain of salt</li>\n<li>Crypto data is very regime-dependent and the public LB score captures a different regime compared to the four folds in my local CV score. For example, one could clearly see how the label distribution is different in the first half of 2023 and  the second half 2023</li>\n</ul>\n<p>Having these two observations, they translated to these actions:</p>\n<ul>\n<li>Make a robust forecast via bagging techniques</li>\n<li>Since we are allowed to make two shots (aka submit two notebooks), then shoot one that aims the LB and one that aims the CV</li>\n</ul>\n<p><strong>Chapter 4. Racing</strong><br>\nI decided to use the classic AK-47, i.e a simple linear regression with forward feature selection. Most of the features had very low correlation with the label so it did not make sense to start with all of them and remove one by one via backward selection. My key racing tips and discoveries are:</p>\n<ul>\n<li>Whenever I added new features (whether one at a time or in groups), I only kept them if they improved correlation by more than the third decimal—otherwise, I considered it just noise and didn’t include them.</li>\n<li>Violating linear regression assumptions isn’t necessarily a problem for forecasting. In practice, accuracy matters more than meeting textbook criteria like p-values. For example, variables with large p-values are often highly correlated with other variables (multicollinearity), making their individual contributions hard to interpret. But including such variables doesn’t always hurt accuracy. If you compare a model with just x1 (y ~ x1) to one with both x1 and a highly correlated x2 (y ~ x1 + x2 where x2 is essentially a linear function of x1), both models can forecast equally well—sometimes the multicollinear model is even better. This happens because many different combinations of coefficients can fit the data similarly well; the model just splits “credit” between the variables.  </li>\n<li>Because of this, I optimized for the highest CV score, not spending too much time on p-values or fixing multicollinearity. The only caveat is that too many multicollinear features can make parameters unstable, and if test-time relationships change, forecasts can suffer. Bagging (averaging) helps reduce this instability.</li>\n<li>Smoothing and removing high-leverage points had the biggest impact on model beta stability and improved both leaderboard and CV scores. In crypto, high-leverage outliers pull the regression line much more than in traditional data, so handling them is crucial.  </li>\n<li>One key feature that improved my score was adding the leverage values like this:</li>\n</ul>\n<pre><code>XtX_inv = np.linalg.pinv(X_design.T @ X_design)\nM = X_design @ XtX_inv\nleverages = np.sum(M * X_design, axis=1)\ndata = data.with_columns(\n        pl.from_numpy(leverages, schema=[]).to_series()\n  )\n</code></pre>\n<p>This feature is the 'leverage' feature and essentially models the fat tails, which in crypto happen often.</p>\n<ul>\n<li>My final linear regression model used 22 features: the 15 most correlated with the label, one leverage feature, and six imbalance features.</li>\n<li>Clipping or winsorizing extreme label values did not improve CV score—in fact, it often made it worse. For example, in one of my experiments, I removed the top 2% most extreme labels and did not include them in evaluation (i.e. the test set). However, I left the extreme values in the train set. The Pearson's correlation on test set improved by 0.015. This suggests that those extreme y-values contain real signal and help guide the model in the right direction; their associated features are actually useful. That is if I could predict accurately the extreme values there was about 0.015 score boost to gain!</li>\n</ul>\n<p><strong>Chapter 5. Late laps, and racing ethics</strong><br>\nThroughout the whole journey I focused exclusively on data exploration, feature engineering and modelling. I refrained myself of peeking or attempting to reverse engineer the problem. Towards the finish, I doubled down on feature engineering — the real spirit of the competition (and makes the race much more fun and exciting). Power features, logs, imbalance measures, clipping, all tried; only the leverage feature and regime-aware engineering stuck. Clipping the most extreme labels actually hurt true test performance: those wild points, I realized, were rich with signal if I wanted to win. </p>\n<p>Final lessons for fellow racers:   </p>\n<ul>\n<li>Simple, robust models, tested on solid, time-series CV can outlast overfit magic tricks on tabular, well structured data with lots of noise.</li>\n<li>Noisy data means moving faster isn’t enough—you need to change direction fast when the regime does.</li>\n</ul>\n<p>The checkered flag may have gone down and the race over, but I already have a keen sense for the next time the starting lights go green (hopefully with a time-series API <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a>)</p>",
      "rawMarkdown": "This is a story of curiosity, learning and experimenting with one main goal — chasing alpha through the fog of crypto noise.\n\n**Chapter 1. Quantifying the (enemy) noise**\nAs we all know, crypto data is notoriously known for having large noise to signal ratio. So the first thing I wanted to check is what exactly noise is in our case. A useful trick I've found is just to shoot the imaginary enemy (crypto label) at random. For example producing 100 random forecast each with of pure random numbers (`np.random.rand`, `np.random.norm`) gave a Pearson's correlations anywhere between  -0.005 to 0.006 - random walk, as expected. This essentially translates to: if your CV score moves at the 3rd decimal, you’re just spinning your wheels\". Only 2nd decimal point jumps are actual progress. \n\n**Chapter 2. Setup a solid CV**\nI decided to use four folds time series CV setup.  Each time I trained on historical data and predicted the remaining future months.  I used the following starting forecasting date for the four folds respectively `FCT_DATES = [\"2023-08-31\", \"2023-09-30\", \"2023-10-31\", \"2023-11-30\"]` Essentially the first fold is trained from \"2023-03-01 to 2023-08-31\" and forecasts the remaining six months from \"2023-08-31 to 2024-02-29\". You get the idea of how the remaining train/test folds work. The key (and most time consumed):  having a modular pipeline allowing for quick iteration — data engineering, feature engineering, and modeling.\n\n**Chapter 3. Getting into the race**\nI've observed there is poor correlation between the public LB and my CV.  For example making a forecast that just takes the value of the most correlated feature (X752) produced 0.09 correlation in CV and only 0.06 in the public LB. Often when I add features in CV and correlation increase, the LB correlation would actually decrease. This meant that:\n- LB data is extremely noisy and results should be taken with a grain of salt\n- Crypto data is very regime-dependent and the public LB score captures a different regime compared to the four folds in my local CV score. For example, one could clearly see how the label distribution is different in the first half of 2023 and  the second half 2023\n\nHaving these two observations, they translated to these actions:\n- Make a robust forecast via bagging techniques\n- Since we are allowed to make two shots (aka submit two notebooks), then shoot one that aims the LB and one that aims the CV\n\n**Chapter 4. Racing**\nI decided to use the classic AK-47, i.e a simple linear regression with forward feature selection. Most of the features had very low correlation with the label so it did not make sense to start with all of them and remove one by one via backward selection. My key racing tips and discoveries are:\n- Whenever I added new features (whether one at a time or in groups), I only kept them if they improved correlation by more than the third decimal—otherwise, I considered it just noise and didn’t include them.\n- Violating linear regression assumptions isn’t necessarily a problem for forecasting. In practice, accuracy matters more than meeting textbook criteria like p-values. For example, variables with large p-values are often highly correlated with other variables (multicollinearity), making their individual contributions hard to interpret. But including such variables doesn’t always hurt accuracy. If you compare a model with just x1 (y ~ x1) to one with both x1 and a highly correlated x2 (y ~ x1 + x2 where x2 is essentially a linear function of x1), both models can forecast equally well—sometimes the multicollinear model is even better. This happens because many different combinations of coefficients can fit the data similarly well; the model just splits “credit” between the variables.  \n- Because of this, I optimized for the highest CV score, not spending too much time on p-values or fixing multicollinearity. The only caveat is that too many multicollinear features can make parameters unstable, and if test-time relationships change, forecasts can suffer. Bagging (averaging) helps reduce this instability.\n- Smoothing and removing high-leverage points had the biggest impact on model beta stability and improved both leaderboard and CV scores. In crypto, high-leverage outliers pull the regression line much more than in traditional data, so handling them is crucial.  \n- One key feature that improved my score was adding the leverage values like this:\n```\nXtX_inv = np.linalg.pinv(X_design.T @ X_design)\nM = X_design @ XtX_inv\nleverages = np.sum(M * X_design, axis=1)\ndata = data.with_columns(\n        pl.from_numpy(leverages, schema=[\"add_leverage_column\"]).to_series()\n  )\n```\nThis feature is the 'leverage' feature and essentially models the fat tails, which in crypto happen often.\n- My final linear regression model used 22 features: the 15 most correlated with the label, one leverage feature, and six imbalance features.\n- Clipping or winsorizing extreme label values did not improve CV score—in fact, it often made it worse. For example, in one of my experiments, I removed the top 2% most extreme labels and did not include them in evaluation (i.e. the test set). However, I left the extreme values in the train set. The Pearson's correlation on test set improved by 0.015. This suggests that those extreme y-values contain real signal and help guide the model in the right direction; their associated features are actually useful. That is if I could predict accurately the extreme values there was about 0.015 score boost to gain!\n     \n**Chapter 5. Late laps, and racing ethics**\nThroughout the whole journey I focused exclusively on data exploration, feature engineering and modelling. I refrained myself of peeking or attempting to reverse engineer the problem. Towards the finish, I doubled down on feature engineering — the real spirit of the competition (and makes the race much more fun and exciting). Power features, logs, imbalance measures, clipping, all tried; only the leverage feature and regime-aware engineering stuck. Clipping the most extreme labels actually hurt true test performance: those wild points, I realized, were rich with signal if I wanted to win. \n\nFinal lessons for fellow racers:   \n\n- Simple, robust models, tested on solid, time-series CV can outlast overfit magic tricks on tabular, well structured data with lots of noise.\n- Noisy data means moving faster isn’t enough—you need to change direction fast when the regime does.\n     \nThe checkered flag may have gone down and the race over, but I already have a keen sense for the next time the starting lights go green (hopefully with a time-series API @drwtrading)\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3262553": "This is a story of curiosity, learning and experimenting with one main goal — chasing alpha through the fog of crypto noise.\n\n**Chapter 1. Quantifying the (enemy) noise**\nAs we all know, crypto data is notoriously known for having large noise to signal ratio. So the first thing I wanted to check is what exactly noise is in our case. A useful trick I've found is just to shoot the imaginary enemy (crypto label) at random. For example producing 100 random forecast each with of pure random numbers (`np.random.rand`, `np.random.norm`) gave a Pearson's correlations anywhere between  -0.005 to 0.006 - random walk, as expected. This essentially translates to: if your CV score moves at the 3rd decimal, you’re just spinning your wheels\". Only 2nd decimal point jumps are actual progress. \n\n**Chapter 2. Setup a solid CV**\nI decided to use four folds time series CV setup.  Each time I trained on historical data and predicted the remaining future months.  I used the following starting forecasting date for the four folds respectively `FCT_DATES = [\"2023-08-31\", \"2023-09-30\", \"2023-10-31\", \"2023-11-30\"]` Essentially the first fold is trained from \"2023-03-01 to 2023-08-31\" and forecasts the remaining six months from \"2023-08-31 to 2024-02-29\". You get the idea of how the remaining train/test folds work. The key (and most time consumed):  having a modular pipeline allowing for quick iteration — data engineering, feature engineering, and modeling.\n\n**Chapter 3. Getting into the race**\nI've observed there is poor correlation between the public LB and my CV.  For example making a forecast that just takes the value of the most correlated feature (X752) produced 0.09 correlation in CV and only 0.06 in the public LB. Often when I add features in CV and correlation increase, the LB correlation would actually decrease. This meant that:\n- LB data is extremely noisy and results should be taken with a grain of salt\n- Crypto data is very regime-dependent and the public LB score captures a different regime compared to the four folds in my local CV score. For example, one could clearly see how the label distribution is different in the first half of 2023 and  the second half 2023\n\nHaving these two observations, they translated to these actions:\n- Make a robust forecast via bagging techniques\n- Since we are allowed to make two shots (aka submit two notebooks), then shoot one that aims the LB and one that aims the CV\n\n**Chapter 4. Racing**\nI decided to use the classic AK-47, i.e a simple linear regression with forward feature selection. Most of the features had very low correlation with the label so it did not make sense to start with all of them and remove one by one via backward selection. My key racing tips and discoveries are:\n- Whenever I added new features (whether one at a time or in groups), I only kept them if they improved correlation by more than the third decimal—otherwise, I considered it just noise and didn’t include them.\n- Violating linear regression assumptions isn’t necessarily a problem for forecasting. In practice, accuracy matters more than meeting textbook criteria like p-values. For example, variables with large p-values are often highly correlated with other variables (multicollinearity), making their individual contributions hard to interpret. But including such variables doesn’t always hurt accuracy. If you compare a model with just x1 (y ~ x1) to one with both x1 and a highly correlated x2 (y ~ x1 + x2 where x2 is essentially a linear function of x1), both models can forecast equally well—sometimes the multicollinear model is even better. This happens because many different combinations of coefficients can fit the data similarly well; the model just splits “credit” between the variables.  \n- Because of this, I optimized for the highest CV score, not spending too much time on p-values or fixing multicollinearity. The only caveat is that too many multicollinear features can make parameters unstable, and if test-time relationships change, forecasts can suffer. Bagging (averaging) helps reduce this instability.\n- Smoothing and removing high-leverage points had the biggest impact on model beta stability and improved both leaderboard and CV scores. In crypto, high-leverage outliers pull the regression line much more than in traditional data, so handling them is crucial.  \n- One key feature that improved my score was adding the leverage values like this:\n```\nXtX_inv = np.linalg.pinv(X_design.T @ X_design)\nM = X_design @ XtX_inv\nleverages = np.sum(M * X_design, axis=1)\ndata = data.with_columns(\n        pl.from_numpy(leverages, schema=[\"add_leverage_column\"]).to_series()\n  )\n```\nThis feature is the 'leverage' feature and essentially models the fat tails, which in crypto happen often.\n- My final linear regression model used 22 features: the 15 most correlated with the label, one leverage feature, and six imbalance features.\n- Clipping or winsorizing extreme label values did not improve CV score—in fact, it often made it worse. For example, in one of my experiments, I removed the top 2% most extreme labels and did not include them in evaluation (i.e. the test set). However, I left the extreme values in the train set. The Pearson's correlation on test set improved by 0.015. This suggests that those extreme y-values contain real signal and help guide the model in the right direction; their associated features are actually useful. That is if I could predict accurately the extreme values there was about 0.015 score boost to gain!\n     \n**Chapter 5. Late laps, and racing ethics**\nThroughout the whole journey I focused exclusively on data exploration, feature engineering and modelling. I refrained myself of peeking or attempting to reverse engineer the problem. Towards the finish, I doubled down on feature engineering — the real spirit of the competition (and makes the race much more fun and exciting). Power features, logs, imbalance measures, clipping, all tried; only the leverage feature and regime-aware engineering stuck. Clipping the most extreme labels actually hurt true test performance: those wild points, I realized, were rich with signal if I wanted to win. \n\nFinal lessons for fellow racers:   \n\n- Simple, robust models, tested on solid, time-series CV can outlast overfit magic tricks on tabular, well structured data with lots of noise.\n- Noisy data means moving faster isn’t enough—you need to change direction fast when the regime does.\n     \nThe checkered flag may have gone down and the race over, but I already have a keen sense for the next time the starting lights go green (hopefully with a time-series API @drwtrading)\n"
  }
}