{
  "id": 599068,
  "title": "DRW solution 1st",
  "url": "/competitions/drw-crypto-market-prediction/discussion/599068",
  "author_name": "A_A",
  "post_date": "2025-08-14T04:47:20.895000",
  "votes": 56,
  "comment_count": 19,
  "views": 0,
  "content": "<h1>Preface:</h1>\n<p>This is my first time sharing my solution on Kaggle. I'll try to be concise, and feel free to share your thoughts or ask if there's anything unclear in my writeup:)</p>\n<h1>Modelling:</h1>\n<p>The setting of this competition is rather simple: regression task on tabular data. Usually, a final ensemble of tree-based models + NN will usually perform the best. And usually one of two will constitute the major chunk, the other is just to provide small improvement in the ensemble stage. For example, in this competition, NN(in particular a 3-layer MLP) is the protagonist(single model 0.124 public and 0.131 private) while XGB plays less, based on my experiments.</p>\n<h1>Cross-validation:</h1>\n<p>Purged group time series split, with 6 groups(roughly 2 months in each group), gap=1(forget about what had happened 2 months before/after testing dataset). I have to say that the CV score is also no near to the public LB(much higher), but the correlation is relatively steady, meaning that an increase in my local CV leads to a consistent increase in public LB.</p>\n<h1>Feature selection&amp;engineering:</h1>\n<p>This is probably the most important part in this competition.</p>\n<ol>\n<li>There are too many primitive features given, and, as everyone might have noticed, they tend to cluster(i.e. having high intra-correlation, threshold=0.6). So what I first did is to find the medoid of each cluster and used them as \"representatives\". Now the number of features comes to around 60. Next I removed features that are absolutely uncorrelated with the target(e.g. correlation &lt;= 1e-4). Now around 40 features are left.</li>\n<li>Run XGB with 6-fold CV(as mentioned earlier), and use SHAP to get the top 20 features for each fold, find the ones that are consistently appearing in each fold's top-20 feature list, take the union, now roughly 30 features left.</li>\n<li>Try linear combinations of the 30 features left, add in the synthesised features, and redo 2..</li>\n<li>\"Recycle\" some features that were previously dropped(c.f. point 3 in \"Something ad hoc\") After finishing this step, the public LB is around 0.10xx.</li>\n<li>Throw the eventual features into AutoEncoder to synthesise 8 new deep features.(AE was very important in my solution, it gave a boost in both local CV and LB)</li>\n</ol>\n<h1>Something ad hoc:</h1>\n<ol>\n<li>Use SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.</li>\n<li>Use 0.6MSE + 0.4Pearson_corr as training loss, and use Pearson_corr as validation metric (Solely using MSE to train MLP leads to poor correlation performance when training MLP, but MSE works well in training XGB, no sure why…anyone has clue?) Generalise MSE to CVaR does not improve CV based on my experiments.</li>\n<li>Everyone may have noticed that your previously well-functioning models were not performing as good after the update of training/testing datasets. This is probably a sign of overfitting. But for me what I found is that no matter how I tried to add in more regularisations to improve robustness, the result gets worse and could never get back to my previous score. This lead to my thinking that some features that are highly correlated with the target in the training set may lose their predictive power as time passes by. So I tried to add back, one-by-one, the features that were previously dropped to see which one gives improvements.</li>\n</ol>\n<h1>Others:</h1>\n<ol>\n<li>I used full training data in my solution(i.e. all rows), but I also noticed that some people uses a selected subsample as training data, which is new to me and very interesting.</li>\n<li>Linear models are powerful. I guess this is partly because the features are powerful and the crypto markets are not so \"twisted\" yet.</li>\n<li>Robustness of the pipeline is very important, particularly when introducing new features derived from black-boxes like AE.</li>\n</ol>\n<h1>Acknowledgement:</h1>\n<p>Thanks to <a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a> for your brilliant EDA, <a href=\"https://www.kaggle.com/taylorsamarel\" target=\"_blank\">@taylorsamarel</a> for your excellent prototyping at the initial stage of the competition, <a href=\"https://www.kaggle.com/shinchen93\" target=\"_blank\">@shinchen93</a> for voicing concerns of all participants through random seed 700. And thanks to <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> for holding this competition. Thanks to all fellow Kagglers for the wonderful discussion and enlightening posts.</p>",
  "messages": [
    {
      "id": 3269195,
      "postDate": "2025-08-14T04:47:20.897Z",
      "content": "<h1>Preface:</h1>\n<p>This is my first time sharing my solution on Kaggle. I'll try to be concise, and feel free to share your thoughts or ask if there's anything unclear in my writeup:)</p>\n<h1>Modelling:</h1>\n<p>The setting of this competition is rather simple: regression task on tabular data. Usually, a final ensemble of tree-based models + NN will usually perform the best. And usually one of two will constitute the major chunk, the other is just to provide small improvement in the ensemble stage. For example, in this competition, NN(in particular a 3-layer MLP) is the protagonist(single model 0.124 public and 0.131 private) while XGB plays less, based on my experiments.</p>\n<h1>Cross-validation:</h1>\n<p>Purged group time series split, with 6 groups(roughly 2 months in each group), gap=1(forget about what had happened 2 months before/after testing dataset). I have to say that the CV score is also no near to the public LB(much higher), but the correlation is relatively steady, meaning that an increase in my local CV leads to a consistent increase in public LB.</p>\n<h1>Feature selection&amp;engineering:</h1>\n<p>This is probably the most important part in this competition.</p>\n<ol>\n<li>There are too many primitive features given, and, as everyone might have noticed, they tend to cluster(i.e. having high intra-correlation, threshold=0.6). So what I first did is to find the medoid of each cluster and used them as \"representatives\". Now the number of features comes to around 60. Next I removed features that are absolutely uncorrelated with the target(e.g. correlation &lt;= 1e-4). Now around 40 features are left.</li>\n<li>Run XGB with 6-fold CV(as mentioned earlier), and use SHAP to get the top 20 features for each fold, find the ones that are consistently appearing in each fold's top-20 feature list, take the union, now roughly 30 features left.</li>\n<li>Try linear combinations of the 30 features left, add in the synthesised features, and redo 2..</li>\n<li>\"Recycle\" some features that were previously dropped(c.f. point 3 in \"Something ad hoc\") After finishing this step, the public LB is around 0.10xx.</li>\n<li>Throw the eventual features into AutoEncoder to synthesise 8 new deep features.(AE was very important in my solution, it gave a boost in both local CV and LB)</li>\n</ol>\n<h1>Something ad hoc:</h1>\n<ol>\n<li>Use SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.</li>\n<li>Use 0.6MSE + 0.4Pearson_corr as training loss, and use Pearson_corr as validation metric (Solely using MSE to train MLP leads to poor correlation performance when training MLP, but MSE works well in training XGB, no sure why…anyone has clue?) Generalise MSE to CVaR does not improve CV based on my experiments.</li>\n<li>Everyone may have noticed that your previously well-functioning models were not performing as good after the update of training/testing datasets. This is probably a sign of overfitting. But for me what I found is that no matter how I tried to add in more regularisations to improve robustness, the result gets worse and could never get back to my previous score. This lead to my thinking that some features that are highly correlated with the target in the training set may lose their predictive power as time passes by. So I tried to add back, one-by-one, the features that were previously dropped to see which one gives improvements.</li>\n</ol>\n<h1>Others:</h1>\n<ol>\n<li>I used full training data in my solution(i.e. all rows), but I also noticed that some people uses a selected subsample as training data, which is new to me and very interesting.</li>\n<li>Linear models are powerful. I guess this is partly because the features are powerful and the crypto markets are not so \"twisted\" yet.</li>\n<li>Robustness of the pipeline is very important, particularly when introducing new features derived from black-boxes like AE.</li>\n</ol>\n<h1>Acknowledgement:</h1>\n<p>Thanks to <a href=\"https://www.kaggle.com/ahsuna123\" target=\"_blank\">@ahsuna123</a> for your brilliant EDA, <a href=\"https://www.kaggle.com/taylorsamarel\" target=\"_blank\">@taylorsamarel</a> for your excellent prototyping at the initial stage of the competition, <a href=\"https://www.kaggle.com/shinchen93\" target=\"_blank\">@shinchen93</a> for voicing concerns of all participants through random seed 700. And thanks to <a href=\"https://www.kaggle.com/drwtrading\" target=\"_blank\">@drwtrading</a> for holding this competition. Thanks to all fellow Kagglers for the wonderful discussion and enlightening posts.</p>",
      "rawMarkdown": "# Preface:\n\nThis is my first time sharing my solution on Kaggle. I'll try to be concise, and feel free to share your thoughts or ask if there's anything unclear in my writeup:)\n\n# Modelling:\n\nThe setting of this competition is rather simple: regression task on tabular data. Usually, a final ensemble of tree-based models + NN will usually perform the best. And usually one of two will constitute the major chunk, the other is just to provide small improvement in the ensemble stage. For example, in this competition, NN(in particular a 3-layer MLP) is the protagonist(single model 0.124 public and 0.131 private) while XGB plays less, based on my experiments.\n\n# Cross-validation:\n\nPurged group time series split, with 6 groups(roughly 2 months in each group), gap=1(forget about what had happened 2 months before/after testing dataset). I have to say that the CV score is also no near to the public LB(much higher), but the correlation is relatively steady, meaning that an increase in my local CV leads to a consistent increase in public LB.\n\n# Feature selection&engineering:\n\nThis is probably the most important part in this competition.\n1.\tThere are too many primitive features given, and, as everyone might have noticed, they tend to cluster(i.e. having high intra-correlation, threshold=0.6). So what I first did is to find the medoid of each cluster and used them as \"representatives\". Now the number of features comes to around 60. Next I removed features that are absolutely uncorrelated with the target(e.g. correlation <= 1e-4). Now around 40 features are left.\n2.\tRun XGB with 6-fold CV(as mentioned earlier), and use SHAP to get the top 20 features for each fold, find the ones that are consistently appearing in each fold's top-20 feature list, take the union, now roughly 30 features left.\n3.\tTry linear combinations of the 30 features left, add in the synthesised features, and redo 2..\n4.\t\"Recycle\" some features that were previously dropped(c.f. point 3 in \"Something ad hoc\") After finishing this step, the public LB is around 0.10xx.\n5.\tThrow the eventual features into AutoEncoder to synthesise 8 new deep features.(AE was very important in my solution, it gave a boost in both local CV and LB)\n\n# Something ad hoc:\n\n1.\tUse SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.\n2.\tUse 0.6MSE + 0.4Pearson_corr as training loss, and use Pearson_corr as validation metric (Solely using MSE to train MLP leads to poor correlation performance when training MLP, but MSE works well in training XGB, no sure why…anyone has clue?) Generalise MSE to CVaR does not improve CV based on my experiments.\n3.\tEveryone may have noticed that your previously well-functioning models were not performing as good after the update of training/testing datasets. This is probably a sign of overfitting. But for me what I found is that no matter how I tried to add in more regularisations to improve robustness, the result gets worse and could never get back to my previous score. This lead to my thinking that some features that are highly correlated with the target in the training set may lose their predictive power as time passes by. So I tried to add back, one-by-one, the features that were previously dropped to see which one gives improvements.\n\n# Others:\n\n1.\tI used full training data in my solution(i.e. all rows), but I also noticed that some people uses a selected subsample as training data, which is new to me and very interesting.\n2.     Linear models are powerful. I guess this is partly because the features are powerful and the crypto markets are not so \"twisted\" yet.\n3.     Robustness of the pipeline is very important, particularly when introducing new features derived from black-boxes like AE.\n  \n# Acknowledgement:\n\nThanks to @ahsuna123 for your brilliant EDA, @taylorsamarel for your excellent prototyping at the initial stage of the competition, @shinchen93 for voicing concerns of all participants through random seed 700. And thanks to @drwtrading for holding this competition. Thanks to all fellow Kagglers for the wonderful discussion and enlightening posts.\n",
      "votes": 56
    },
    {
      "id": 3295366,
      "postDate": "2025-09-28T14:40:19.190Z",
      "content": "<p>Thanks a lot for sharing your solution! I’m just starting out, but your explanation made some key ideas much clearer. I’ll definitely try using the purged group time series split and feature clustering in my own work. Really appreciate you taking the time to break it down.</p>",
      "rawMarkdown": "Thanks a lot for sharing your solution! I’m just starting out, but your explanation made some key ideas much clearer. I’ll definitely try using the purged group time series split and feature clustering in my own work. Really appreciate you taking the time to break it down."
    },
    {
      "id": 3281308,
      "postDate": "2025-09-04T03:23:29.603Z",
      "content": "<p>Hi Tony,</p>\n<p>On the second point of your feature selection step, I am confused about why you taking the union of the top 20 features from each fold. I cannot see how the union could identify \"the ones that are consistently appearing in each fold's top-20 feature list\". Could you please shed some light on this? Thanks a lot!</p>",
      "rawMarkdown": "Hi Tony,\n\nOn the second point of your feature selection step, I am confused about why you taking the union of the top 20 features from each fold. I cannot see how the union could identify \"the ones that are consistently appearing in each fold's top-20 feature list\". Could you please shed some light on this? Thanks a lot!",
      "replies": [
        {
          "id": 3281318,
          "postDate": "2025-09-04T03:41:28.930Z",
          "content": "<p>The idea of this step is to find features that are \"important\" to all folds, or features that are robust enough. For example, if you found X_1 ranked 15th in the 1st fold while having little significance in the rest 5 folds, you probably want to drop it. (Again, hyper-parameters used here were bit arbitrary)</p>",
          "rawMarkdown": "The idea of this step is to find features that are \"important\" to all folds, or features that are robust enough. For example, if you found X_1 ranked 15th in the 1st fold while having little significance in the rest 5 folds, you probably want to drop it. (Again, hyper-parameters used here were bit arbitrary)",
          "replies": [
            {
              "id": 3281325,
              "postDate": "2025-09-04T03:54:29.357Z",
              "content": "<p>If I understand correctly, for each of the 6 models trained in the CV, you were taking the intersection of top 20 features from each fold, and this results in 6 lists of features, one for each model. Lastly, you took the union of these 6 lists and results in roughly 30 features left. Does this sound like right?</p>",
              "rawMarkdown": "If I understand correctly, for each of the 6 models trained in the CV, you were taking the intersection of top 20 features from each fold, and this results in 6 lists of features, one for each model. Lastly, you took the union of these 6 lists and results in roughly 30 features left. Does this sound like right?"
            },
            {
              "id": 3281381,
              "postDate": "2025-09-04T07:36:36.847Z",
              "content": "<p>That's correct</p>",
              "rawMarkdown": "That's correct"
            }
          ]
        }
      ]
    },
    {
      "id": 3275469,
      "postDate": "2025-08-26T15:22:54.307Z",
      "content": "<p>Amazing performance….</p>",
      "rawMarkdown": "Amazing performance...."
    },
    {
      "id": 3274546,
      "postDate": "2025-08-25T02:02:44.210Z",
      "content": "<p>Hi Tony,</p>\n<p>I have a question regarding your feature selection step. When you mention reducing the features to around 60 by clustering, could you clarify:</p>\n<blockquote>\n  <p>they tend to cluster(i.e. having high intra-correlation, threshold=0.6)</p>\n</blockquote>\n<ul>\n<li><p>Which clustering method or model did you use to group the features?</p></li>\n<li><p>Was the number of clusters (≈60) determined by setting a specific hyperparameter, or did it naturally result from the correlation threshold?</p></li>\n<li><p>When selecting the medoid as the representative feature of each cluster, how exactly was it chosen? Was it simply the central point of the cluster or based on another criterion?</p></li>\n</ul>\n<p>Thank you very much for your time and for sharing this insightful solution.</p>",
      "rawMarkdown": "Hi Tony,\n\nI have a question regarding your feature selection step. When you mention reducing the features to around 60 by clustering, could you clarify:\n\n>they tend to cluster(i.e. having high intra-correlation, threshold=0.6)\n\n- Which clustering method or model did you use to group the features?\n\n- Was the number of clusters (≈60) determined by setting a specific hyperparameter, or did it naturally result from the correlation threshold?\n\n- When selecting the medoid as the representative feature of each cluster, how exactly was it chosen? Was it simply the central point of the cluster or based on another criterion?\n\nThank you very much for your time and for sharing this insightful solution.",
      "replies": [
        {
          "id": 3275406,
          "postDate": "2025-08-26T13:00:40.367Z",
          "content": "<ol>\n<li>I just used 1-abs(corr) as a distance metric and solved a combinatorial optimization problem..(but I guess directly using KMeans or other packages also work out)</li>\n<li>60 is a result of the threshold 0.6(you could get more/less clusters by setting threshold lower/higher)</li>\n<li>Medoid is like an extension of the concept of median, here it represents the element with highest sum of abs(correlation) with other elements in a cluster.(i.e. the most expensive point to be replicated from other points)</li>\n</ol>\n<p>Hope this clarifies.</p>",
          "rawMarkdown": "1. I just used 1-abs(corr) as a distance metric and solved a combinatorial optimization problem..(but I guess directly using KMeans or other packages also work out)\n2. 60 is a result of the threshold 0.6(you could get more/less clusters by setting threshold lower/higher)\n3. Medoid is like an extension of the concept of median, here it represents the element with highest sum of abs(correlation) with other elements in a cluster.(i.e. the most expensive point to be replicated from other points)\n\nHope this clarifies.",
          "votes": 2,
          "replies": [
            {
              "id": 3276874,
              "postDate": "2025-08-27T00:34:02.500Z",
              "content": "<p>Hi Tony,</p>\n<p>I really appreciate your clarifications. Since I just started exploring Kaggle competitions, there are still many things that are not very clear to me, and I’m very grateful for your patience in answering my questions.</p>\n<p>I do have one more question: I noticed you chose an MLP instead of a time-series model. Because my first thought was that, since this is a time-series competition, a time-series model might be more natural. Did you try both approaches and find that MLP performed better in practice? Or was there another reason behind preferring MLP?</p>\n<p>Thanks again for your help!</p>",
              "rawMarkdown": "Hi Tony,\n\nI really appreciate your clarifications. Since I just started exploring Kaggle competitions, there are still many things that are not very clear to me, and I’m very grateful for your patience in answering my questions.\n\nI do have one more question: I noticed you chose an MLP instead of a time-series model. Because my first thought was that, since this is a time-series competition, a time-series model might be more natural. Did you try both approaches and find that MLP performed better in practice? Or was there another reason behind preferring MLP?\n\nThanks again for your help!"
            },
            {
              "id": 3276924,
              "postDate": "2025-08-27T02:53:40.487Z",
              "content": "<p>This is not a time-series competition.</p>",
              "rawMarkdown": "This is not a time-series competition."
            }
          ]
        }
      ]
    },
    {
      "id": 3273948,
      "postDate": "2025-08-23T16:55:31.647Z",
      "content": "<p>Excellent work! Really cool to see the MLP architecture perform so well!</p>",
      "rawMarkdown": "Excellent work! Really cool to see the MLP architecture perform so well!"
    },
    {
      "id": 3269414,
      "postDate": "2025-08-14T13:04:54.503Z",
      "content": "<p>Inspired work! I got confused on point 3 for linear combinations of the 30 features left. What's the method you are using to generate new feature combinations? What's the feature dimension in the end?</p>",
      "rawMarkdown": "Inspired work! I got confused on point 3 for linear combinations of the 30 features left. What's the method you are using to generate new feature combinations? What's the feature dimension in the end?",
      "replies": [
        {
          "id": 3269495,
          "postDate": "2025-08-14T15:05:43.940Z",
          "content": "<p>I guess I didn't state this one correctly. In fact not just linear combinations, I tried 2nd order combinations like X+-*/Y, max min(X,Y), and 3rd order combinations like XYZ, (X+Y)/Z, max(min(X,Y),Z) etc., where X,Y,Z are 1d feature vectors from the 30 features(i.e. one column), and the features obtained are thus all 1 dimensional. Hope this clarifies.</p>",
          "rawMarkdown": "I guess I didn't state this one correctly. In fact not just linear combinations, I tried 2nd order combinations like X+-*/Y, max min(X,Y), and 3rd order combinations like XYZ, (X+Y)/Z, max(min(X,Y),Z) etc., where X,Y,Z are 1d feature vectors from the 30 features(i.e. one column), and the features obtained are thus all 1 dimensional. Hope this clarifies.",
          "replies": [
            {
              "id": 3269753,
              "postDate": "2025-08-15T02:04:25.317Z",
              "content": "<p>Thanks for clarification. If my undestand is correct, you are doing combinations among filtered features, like symbolic regression, and deriving those features with higher correlation with target.</p>",
              "rawMarkdown": "Thanks for clarification. If my undestand is correct, you are doing combinations among filtered features, like symbolic regression, and deriving those features with higher correlation with target."
            },
            {
              "id": 3270010,
              "postDate": "2025-08-15T14:22:01.060Z",
              "content": "<p>Yes. Something like that.</p>",
              "rawMarkdown": "Yes. Something like that."
            }
          ]
        }
      ]
    },
    {
      "id": 3269388,
      "postDate": "2025-08-14T12:14:12.193Z",
      "content": "<p>Impressive performance! </p>\n<blockquote>\n  <p>Purged group time series split, with 6 groups</p>\n</blockquote>\n<p>Could you share your CV results and the score for each fold? I used Purged K-Fold with an embargo, but it seems that  Purged Group Time Series Split offers more robust and reliable evaluation.</p>\n<blockquote>\n  <p>Use SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.</p>\n</blockquote>\n<p>SGD is known to generalize better than Adam/AdamW. In this competition, I experimented with Muon, which outperformed all other optimizers and delivered strong generalization on both the public and private leaderboards. After the competition ended, I also tried an MLP, and once again, Muon impressed me - it truly surprised me!</p>\n<blockquote>\n  <p>Solely using MSE to train MLP leads to poor correlation performance when training MLP</p>\n</blockquote>\n<p>One intuition is that, in deep learning, the best loss is often the metric itself 🙂 - you’re directly optimizing what you care about. Since DNNs are universal approximators, they can adapt to fit that metric closely.</p>\n<p>Btw, do you intend to share your training code? </p>",
      "rawMarkdown": "Impressive performance! \n\n>Purged group time series split, with 6 groups\n\nCould you share your CV results and the score for each fold? I used Purged K-Fold with an embargo, but it seems that  Purged Group Time Series Split offers more robust and reliable evaluation.\n\n>Use SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.\n\nSGD is known to generalize better than Adam/AdamW. In this competition, I experimented with Muon, which outperformed all other optimizers and delivered strong generalization on both the public and private leaderboards. After the competition ended, I also tried an MLP, and once again, Muon impressed me - it truly surprised me!\n\n> Solely using MSE to train MLP leads to poor correlation performance when training MLP\n\nOne intuition is that, in deep learning, the best loss is often the metric itself 🙂 - you’re directly optimizing what you care about. Since DNNs are universal approximators, they can adapt to fit that metric closely.\n\nBtw, do you intend to share your training code? ",
      "replies": [
        {
          "id": 3269484,
          "postDate": "2025-08-14T14:58:13.353Z",
          "content": "<p>The pearson_r of local CVs are roughly 0.1961/0.1598/0.0338/0.1420/0.1457/0.1114, and it seems that excluding the \"bad\" fold doesn't affect much based on my experiment. And thanks for the insightful comments! I also thought about using Muon during the training stage but failed to do so due to time constraints, perhaps I should learn and try more about it! Currently I'm not considering to make my code public as it is completely a mess 😅 </p>",
          "rawMarkdown": "The pearson_r of local CVs are roughly 0.1961/0.1598/0.0338/0.1420/0.1457/0.1114, and it seems that excluding the \"bad\" fold doesn't affect much based on my experiment. And thanks for the insightful comments! I also thought about using Muon during the training stage but failed to do so due to time constraints, perhaps I should learn and try more about it! Currently I'm not considering to make my code public as it is completely a mess 😅 ",
          "votes": 1,
          "replies": [
            {
              "id": 3269491,
              "postDate": "2025-08-14T15:01:52.360Z",
              "content": "<p>Thanks! Hope you could consider making it public! Most people might benefit from your code and good practices!</p>",
              "rawMarkdown": "Thanks! Hope you could consider making it public! Most people might benefit from your code and good practices!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3269400,
      "postDate": "2025-08-14T12:46:43.530Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3295366,
      "author_name": "n00b0dyy",
      "author_url": "",
      "post_date": "2025-09-28T14:40:19.190000",
      "content": "<p>Thanks a lot for sharing your solution! I’m just starting out, but your explanation made some key ideas much clearer. I’ll definitely try using the purged group time series split and feature clustering in my own work. Really appreciate you taking the time to break it down.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3281308,
      "author_name": "Gabriella Chaos",
      "author_url": "",
      "post_date": "2025-09-04T03:23:29.603000",
      "content": "<p>Hi Tony,</p>\n<p>On the second point of your feature selection step, I am confused about why you taking the union of the top 20 features from each fold. I cannot see how the union could identify \"the ones that are consistently appearing in each fold's top-20 feature list\". Could you please shed some light on this? Thanks a lot!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3281318,
          "author_name": "A_A",
          "author_url": "",
          "post_date": "2025-09-04T03:41:28.930000",
          "content": "<p>The idea of this step is to find features that are \"important\" to all folds, or features that are robust enough. For example, if you found X_1 ranked 15th in the 1st fold while having little significance in the rest 5 folds, you probably want to drop it. (Again, hyper-parameters used here were bit arbitrary)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3281325,
              "author_name": "Gabriella Chaos",
              "author_url": "",
              "post_date": "2025-09-04T03:54:29.357000",
              "content": "<p>If I understand correctly, for each of the 6 models trained in the CV, you were taking the intersection of top 20 features from each fold, and this results in 6 lists of features, one for each model. Lastly, you took the union of these 6 lists and results in roughly 30 features left. Does this sound like right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3281381,
              "author_name": "A_A",
              "author_url": "",
              "post_date": "2025-09-04T07:36:36.847000",
              "content": "<p>That's correct</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3275469,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-08-26T15:22:54.307000",
      "content": "<p>Amazing performance….</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3274546,
      "author_name": "JJ(Eric) Hu",
      "author_url": "",
      "post_date": "2025-08-25T02:02:44.210000",
      "content": "<p>Hi Tony,</p>\n<p>I have a question regarding your feature selection step. When you mention reducing the features to around 60 by clustering, could you clarify:</p>\n<blockquote>\n  <p>they tend to cluster(i.e. having high intra-correlation, threshold=0.6)</p>\n</blockquote>\n<ul>\n<li><p>Which clustering method or model did you use to group the features?</p></li>\n<li><p>Was the number of clusters (≈60) determined by setting a specific hyperparameter, or did it naturally result from the correlation threshold?</p></li>\n<li><p>When selecting the medoid as the representative feature of each cluster, how exactly was it chosen? Was it simply the central point of the cluster or based on another criterion?</p></li>\n</ul>\n<p>Thank you very much for your time and for sharing this insightful solution.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3275406,
          "author_name": "A_A",
          "author_url": "",
          "post_date": "2025-08-26T13:00:40.367000",
          "content": "<ol>\n<li>I just used 1-abs(corr) as a distance metric and solved a combinatorial optimization problem..(but I guess directly using KMeans or other packages also work out)</li>\n<li>60 is a result of the threshold 0.6(you could get more/less clusters by setting threshold lower/higher)</li>\n<li>Medoid is like an extension of the concept of median, here it represents the element with highest sum of abs(correlation) with other elements in a cluster.(i.e. the most expensive point to be replicated from other points)</li>\n</ol>\n<p>Hope this clarifies.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3276874,
              "author_name": "JJ(Eric) Hu",
              "author_url": "",
              "post_date": "2025-08-27T00:34:02.500000",
              "content": "<p>Hi Tony,</p>\n<p>I really appreciate your clarifications. Since I just started exploring Kaggle competitions, there are still many things that are not very clear to me, and I’m very grateful for your patience in answering my questions.</p>\n<p>I do have one more question: I noticed you chose an MLP instead of a time-series model. Because my first thought was that, since this is a time-series competition, a time-series model might be more natural. Did you try both approaches and find that MLP performed better in practice? Or was there another reason behind preferring MLP?</p>\n<p>Thanks again for your help!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3276924,
              "author_name": "A_A",
              "author_url": "",
              "post_date": "2025-08-27T02:53:40.487000",
              "content": "<p>This is not a time-series competition.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3273948,
      "author_name": "Taylor S. Amarel",
      "author_url": "",
      "post_date": "2025-08-23T16:55:31.647000",
      "content": "<p>Excellent work! Really cool to see the MLP architecture perform so well!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3269414,
      "author_name": "Henry",
      "author_url": "",
      "post_date": "2025-08-14T13:04:54.503000",
      "content": "<p>Inspired work! I got confused on point 3 for linear combinations of the 30 features left. What's the method you are using to generate new feature combinations? What's the feature dimension in the end?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3269495,
          "author_name": "A_A",
          "author_url": "",
          "post_date": "2025-08-14T15:05:43.940000",
          "content": "<p>I guess I didn't state this one correctly. In fact not just linear combinations, I tried 2nd order combinations like X+-*/Y, max min(X,Y), and 3rd order combinations like XYZ, (X+Y)/Z, max(min(X,Y),Z) etc., where X,Y,Z are 1d feature vectors from the 30 features(i.e. one column), and the features obtained are thus all 1 dimensional. Hope this clarifies.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3269753,
              "author_name": "Henry",
              "author_url": "",
              "post_date": "2025-08-15T02:04:25.317000",
              "content": "<p>Thanks for clarification. If my undestand is correct, you are doing combinations among filtered features, like symbolic regression, and deriving those features with higher correlation with target.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3270010,
              "author_name": "A_A",
              "author_url": "",
              "post_date": "2025-08-15T14:22:01.060000",
              "content": "<p>Yes. Something like that.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3269388,
      "author_name": "ducnh279",
      "author_url": "",
      "post_date": "2025-08-14T12:14:12.193000",
      "content": "<p>Impressive performance! </p>\n<blockquote>\n  <p>Purged group time series split, with 6 groups</p>\n</blockquote>\n<p>Could you share your CV results and the score for each fold? I used Purged K-Fold with an embargo, but it seems that  Purged Group Time Series Split offers more robust and reliable evaluation.</p>\n<blockquote>\n  <p>Use SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.</p>\n</blockquote>\n<p>SGD is known to generalize better than Adam/AdamW. In this competition, I experimented with Muon, which outperformed all other optimizers and delivered strong generalization on both the public and private leaderboards. After the competition ended, I also tried an MLP, and once again, Muon impressed me - it truly surprised me!</p>\n<blockquote>\n  <p>Solely using MSE to train MLP leads to poor correlation performance when training MLP</p>\n</blockquote>\n<p>One intuition is that, in deep learning, the best loss is often the metric itself 🙂 - you’re directly optimizing what you care about. Since DNNs are universal approximators, they can adapt to fit that metric closely.</p>\n<p>Btw, do you intend to share your training code? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3269484,
          "author_name": "A_A",
          "author_url": "",
          "post_date": "2025-08-14T14:58:13.353000",
          "content": "<p>The pearson_r of local CVs are roughly 0.1961/0.1598/0.0338/0.1420/0.1457/0.1114, and it seems that excluding the \"bad\" fold doesn't affect much based on my experiment. And thanks for the insightful comments! I also thought about using Muon during the training stage but failed to do so due to time constraints, perhaps I should learn and try more about it! Currently I'm not considering to make my code public as it is completely a mess 😅 </p>",
          "votes": 1,
          "replies": [
            {
              "id": 3269491,
              "author_name": "ducnh279",
              "author_url": "",
              "post_date": "2025-08-14T15:01:52.360000",
              "content": "<p>Thanks! Hope you could consider making it public! Most people might benefit from your code and good practices!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3269400,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-08-14T12:46:43.530000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3269195": "# Preface:\n\nThis is my first time sharing my solution on Kaggle. I'll try to be concise, and feel free to share your thoughts or ask if there's anything unclear in my writeup:)\n\n# Modelling:\n\nThe setting of this competition is rather simple: regression task on tabular data. Usually, a final ensemble of tree-based models + NN will usually perform the best. And usually one of two will constitute the major chunk, the other is just to provide small improvement in the ensemble stage. For example, in this competition, NN(in particular a 3-layer MLP) is the protagonist(single model 0.124 public and 0.131 private) while XGB plays less, based on my experiments.\n\n# Cross-validation:\n\nPurged group time series split, with 6 groups(roughly 2 months in each group), gap=1(forget about what had happened 2 months before/after testing dataset). I have to say that the CV score is also no near to the public LB(much higher), but the correlation is relatively steady, meaning that an increase in my local CV leads to a consistent increase in public LB.\n\n# Feature selection&engineering:\n\nThis is probably the most important part in this competition.\n1.\tThere are too many primitive features given, and, as everyone might have noticed, they tend to cluster(i.e. having high intra-correlation, threshold=0.6). So what I first did is to find the medoid of each cluster and used them as \"representatives\". Now the number of features comes to around 60. Next I removed features that are absolutely uncorrelated with the target(e.g. correlation <= 1e-4). Now around 40 features are left.\n2.\tRun XGB with 6-fold CV(as mentioned earlier), and use SHAP to get the top 20 features for each fold, find the ones that are consistently appearing in each fold's top-20 feature list, take the union, now roughly 30 features left.\n3.\tTry linear combinations of the 30 features left, add in the synthesised features, and redo 2..\n4.\t\"Recycle\" some features that were previously dropped(c.f. point 3 in \"Something ad hoc\") After finishing this step, the public LB is around 0.10xx.\n5.\tThrow the eventual features into AutoEncoder to synthesise 8 new deep features.(AE was very important in my solution, it gave a boost in both local CV and LB)\n\n# Something ad hoc:\n\n1.\tUse SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.\n2.\tUse 0.6MSE + 0.4Pearson_corr as training loss, and use Pearson_corr as validation metric (Solely using MSE to train MLP leads to poor correlation performance when training MLP, but MSE works well in training XGB, no sure why…anyone has clue?) Generalise MSE to CVaR does not improve CV based on my experiments.\n3.\tEveryone may have noticed that your previously well-functioning models were not performing as good after the update of training/testing datasets. This is probably a sign of overfitting. But for me what I found is that no matter how I tried to add in more regularisations to improve robustness, the result gets worse and could never get back to my previous score. This lead to my thinking that some features that are highly correlated with the target in the training set may lose their predictive power as time passes by. So I tried to add back, one-by-one, the features that were previously dropped to see which one gives improvements.\n\n# Others:\n\n1.\tI used full training data in my solution(i.e. all rows), but I also noticed that some people uses a selected subsample as training data, which is new to me and very interesting.\n2.     Linear models are powerful. I guess this is partly because the features are powerful and the crypto markets are not so \"twisted\" yet.\n3.     Robustness of the pipeline is very important, particularly when introducing new features derived from black-boxes like AE.\n  \n# Acknowledgement:\n\nThanks to @ahsuna123 for your brilliant EDA, @taylorsamarel for your excellent prototyping at the initial stage of the competition, @shinchen93 for voicing concerns of all participants through random seed 700. And thanks to @drwtrading for holding this competition. Thanks to all fellow Kagglers for the wonderful discussion and enlightening posts.\n",
    "3295366": "Thanks a lot for sharing your solution! I’m just starting out, but your explanation made some key ideas much clearer. I’ll definitely try using the purged group time series split and feature clustering in my own work. Really appreciate you taking the time to break it down.",
    "3281308": "Hi Tony,\n\nOn the second point of your feature selection step, I am confused about why you taking the union of the top 20 features from each fold. I cannot see how the union could identify \"the ones that are consistently appearing in each fold's top-20 feature list\". Could you please shed some light on this? Thanks a lot!",
    "3275469": "Amazing performance....",
    "3274546": "Hi Tony,\n\nI have a question regarding your feature selection step. When you mention reducing the features to around 60 by clustering, could you clarify:\n\n>they tend to cluster(i.e. having high intra-correlation, threshold=0.6)\n\n- Which clustering method or model did you use to group the features?\n\n- Was the number of clusters (≈60) determined by setting a specific hyperparameter, or did it naturally result from the correlation threshold?\n\n- When selecting the medoid as the representative feature of each cluster, how exactly was it chosen? Was it simply the central point of the cluster or based on another criterion?\n\nThank you very much for your time and for sharing this insightful solution.",
    "3273948": "Excellent work! Really cool to see the MLP architecture perform so well!",
    "3269414": "Inspired work! I got confused on point 3 for linear combinations of the 30 features left. What's the method you are using to generate new feature combinations? What's the feature dimension in the end?",
    "3269388": "Impressive performance! \n\n>Purged group time series split, with 6 groups\n\nCould you share your CV results and the score for each fold? I used Purged K-Fold with an embargo, but it seems that  Purged Group Time Series Split offers more robust and reliable evaluation.\n\n>Use SGD instead of Adam, AdamW to train MLP. For me the later 2 optimisers work poorly.\n\nSGD is known to generalize better than Adam/AdamW. In this competition, I experimented with Muon, which outperformed all other optimizers and delivered strong generalization on both the public and private leaderboards. After the competition ended, I also tried an MLP, and once again, Muon impressed me - it truly surprised me!\n\n> Solely using MSE to train MLP leads to poor correlation performance when training MLP\n\nOne intuition is that, in deep learning, the best loss is often the metric itself 🙂 - you’re directly optimizing what you care about. Since DNNs are universal approximators, they can adapt to fit that metric closely.\n\nBtw, do you intend to share your training code? ",
    "3269400": ""
  }
}