{
  "id": 596581,
  "title": " Emphasizing Data Quality over Model Complexity (2nd & 25th Place Solutions)",
  "url": "/competitions/drw-crypto-market-prediction/writeups/emphasizing-data-quality-over-model-complexity-2nd",
  "author_name": "",
  "post_date": "2025-08-04T13:00:05.270Z",
  "votes": 12,
  "comment_count": 9,
  "views": 0,
  "content": "<p>My approach focused more on <strong>data preparation</strong> than on overly complex modeling. I started by engineering a time-based feature, <code>Y_M_D_H</code>, extracted from the timestamp column, which had a 1-minute resolution.<br>\nUsing this feature, I grouped the data into <strong>hourly intervals</strong> and applied <strong>clustering techniques</strong> on selected features to detect similarity patterns across rows.<br>\nThen, through an <strong>iterative filtering process</strong>, I removed or retained specific rows based on their impact on the public leaderboard score. For instance, if removing a cluster of rows (belonging to a particular <code>Y_M_D_H</code> group) improved validation performance, I excluded them permanently. If performance dropped, I kept them. This process helped me build a <strong>cleaned, optimized training set</strong> for classical models.<br>\nThe goal wasn't to create a perfect preprocessing pipeline, but to show that <strong>clean, well-structured training data can outperform complex models trained on noisy or unfiltered data</strong>.<br>\nInterestingly:</p>\n<ul>\n<li>For my <strong>25th place solution</strong>, I used a <code>1-hour</code> timestamp.</li>\n<li>For my <strong>2nd place private leaderboard solution</strong>, I used a more granular <code>1-minute</code> timestamp, forming a <code>Y_M_D_H_M</code> feature.<br>\nIn both cases, I used <strong>the same model and features</strong> — only the <strong>data preparation strategy differed</strong>.</li>\n<li>📄 <a href=\"https://www.kaggle.com/code/youneseloiarm/drw-2th-place-in-the-private-lb\" target=\"_blank\">2nd Place Solution (Private LB)</a></li>\n<li>📄 <a href=\"https://www.kaggle.com/code/youneseloiarm/drw-25th-place-in-the-private-lb\" target=\"_blank\">25th Place Solution (Private LB)</a></li>\n</ul>",
  "messages": [
    {
      "id": "3262903",
      "postDate": "08/04/2025 12:56:21",
      "content": "<p>My approach focused more on <strong>data preparation</strong> than on overly complex modeling. I started by engineering a time-based feature, <code>Y_M_D_H</code>, extracted from the timestamp column, which had a 1-minute resolution.<br>\nUsing this feature, I grouped the data into <strong>hourly intervals</strong> and applied <strong>clustering techniques</strong> on selected features to detect similarity patterns across rows.<br>\nThen, through an <strong>iterative filtering process</strong>, I removed or retained specific rows based on their impact on the public leaderboard score. For instance, if removing a cluster of rows (belonging to a particular <code>Y_M_D_H</code> group) improved validation performance, I excluded them permanently. If performance dropped, I kept them. This process helped me build a <strong>cleaned, optimized training set</strong> for classical models.<br>\nThe goal wasn't to create a perfect preprocessing pipeline, but to show that <strong>clean, well-structured training data can outperform complex models trained on noisy or unfiltered data</strong>.<br>\nInterestingly:</p>\n<ul>\n<li>For my <strong>25th place solution</strong>, I used a <code>1-hour</code> timestamp.</li>\n<li>For my <strong>2nd place private leaderboard solution</strong>, I used a more granular <code>1-minute</code> timestamp, forming a <code>Y_M_D_H_M</code> feature.<br>\nIn both cases, I used <strong>the same model and features</strong> — only the <strong>data preparation strategy differed</strong>.</li>\n<li>📄 <a href=\"https://www.kaggle.com/code/youneseloiarm/drw-2th-place-in-the-private-lb\" target=\"_blank\">2nd Place Solution (Private LB)</a></li>\n<li>📄 <a href=\"https://www.kaggle.com/code/youneseloiarm/drw-25th-place-in-the-private-lb\" target=\"_blank\">25th Place Solution (Private LB)</a></li>\n</ul>",
      "rawMarkdown": "My approach focused more on **data preparation** than on overly complex modeling. I started by engineering a time-based feature, `Y_M_D_H`, extracted from the timestamp column, which had a 1-minute resolution.\n\nUsing this feature, I grouped the data into **hourly intervals** and applied **clustering techniques** on selected features to detect similarity patterns across rows.\n\nThen, through an **iterative filtering process**, I removed or retained specific rows based on their impact on the public leaderboard score. For instance, if removing a cluster of rows (belonging to a particular `Y_M_D_H` group) improved validation performance, I excluded them permanently. If performance dropped, I kept them. This process helped me build a **cleaned, optimized training set** for classical models.\n\nThe goal wasn't to create a perfect preprocessing pipeline, but to show that **clean, well-structured training data can outperform complex models trained on noisy or unfiltered data**.\n\nInterestingly:\n\n* For my **25th place solution**, I used a `1-hour` timestamp.\n* For my **2nd place private leaderboard solution**, I used a more granular `1-minute` timestamp, forming a `Y_M_D_H_M` feature.\n\nIn both cases, I used **the same model and features** — only the **data preparation strategy differed**.\n\n* 📄 [2nd Place Solution (Private LB)](https://www.kaggle.com/code/youneseloiarm/drw-2th-place-in-the-private-lb)\n* 📄 [25th Place Solution (Private LB)](https://www.kaggle.com/code/youneseloiarm/drw-25th-place-in-the-private-lb)",
      "votes": null
    },
    {
      "id": "3262911",
      "postDate": "08/04/2025 13:14:15",
      "content": "<p>Hi, thanks for sharing your solution! </p>\n<p>It seems you did not perform cross-validation? if yes, what is your cross-validation strategy and its score?</p>",
      "rawMarkdown": "Hi, thanks for sharing your solution! \n\nIt seems you did not perform cross-validation? if yes, what is your cross-validation strategy and its score?",
      "votes": null
    },
    {
      "id": "3262917",
      "postDate": "08/04/2025 13:24:04",
      "content": "<p>Linear models are among the oldest modeling techniques, commonly used in theoretical economic models and economic research in general. One key advantage is that they tend not to overfit as easily as many machine learning or deep learning models.</p>\n<p>However, you need to be careful with the features you feed into the model. If the features are meaningful and well-engineered, the linear model will likely perform well—it's that simple.</p>\n<p>As I mentioned earlier, the public leaderboard reflects 49% of the data, so I focused my optimization efforts accordingly.</p>",
      "rawMarkdown": "Linear models are among the oldest modeling techniques, commonly used in theoretical economic models and economic research in general. One key advantage is that they tend not to overfit as easily as many machine learning or deep learning models.\n\nHowever, you need to be careful with the features you feed into the model. If the features are meaningful and well-engineered, the linear model will likely perform well—it's that simple.\n\nAs I mentioned earlier, the public leaderboard reflects 49% of the data, so I focused my optimization efforts accordingly.",
      "votes": null
    },
    {
      "id": "3262921",
      "postDate": "08/04/2025 13:32:26",
      "content": "<p>I’m a bit confused about your approach and your reply, but anyway - congrats on finishing 25th! 🎉</p>",
      "rawMarkdown": "I’m a bit confused about your approach and your reply, but anyway - congrats on finishing 25th! 🎉",
      "votes": null
    },
    {
      "id": "3264118",
      "postDate": "08/06/2025 07:28:39",
      "content": "<p>Thank you for sharing your nice solution. May I ask how did you come up with the idea of removing certain group of row to make a better train set.</p>",
      "rawMarkdown": "Thank you for sharing your nice solution. May I ask how did you come up with the idea of removing certain group of row to make a better train set.",
      "votes": null
    },
    {
      "id": "3264311",
      "postDate": "08/06/2025 13:03:40",
      "content": "<p>Thanks for your kind words! Yes, the idea of removing certain groups of rows actually came from my experience working with financial time series data, where row-level filtering often helps improve generalization and stability. In those settings, some rows introduce noise or shift distribution in a way that harms model performance. So I applied the same intuition here. Curious—did you try something similar in your approach?</p>",
      "rawMarkdown": "Thanks for your kind words! Yes, the idea of removing certain groups of rows actually came from my experience working with financial time series data, where row-level filtering often helps improve generalization and stability. In those settings, some rows introduce noise or shift distribution in a way that harms model performance. So I applied the same intuition here. Curious—did you try something similar in your approach?",
      "votes": null
    },
    {
      "id": "3264692",
      "postDate": "08/06/2025 20:19:17",
      "content": "<p>Wow that was interesting, I have only thought maybe removing outlier or sth simpler. For me I think having an ensemble of model where each model is trained on certain fixed sized window made my solution a little bit more robust. Other than that I mainly use XGBoost and MLP (which I found that less hidden neurons =&gt; more robust =&gt; better score in this scenario) on selected features. I certainly got a lot of luck when the private score was released. Anw, thank you for sharing!</p>",
      "rawMarkdown": "Wow that was interesting, I have only thought maybe removing outlier or sth simpler. For me I think having an ensemble of model where each model is trained on certain fixed sized window made my solution a little bit more robust. Other than that I mainly use XGBoost and MLP (which I found that less hidden neurons => more robust => better score in this scenario) on selected features. I certainly got a lot of luck when the private score was released. Anw, thank you for sharing!",
      "votes": null
    },
    {
      "id": "3264849",
      "postDate": "08/06/2025 22:31:19",
      "content": "<p>Hi, thanks for the writeup.</p>\n<p>Can you please go into more detail on how you selected the features to apply the clustering algorithm on? I think this is a vital part that will be helpful to fully understand the data quality argument here.</p>",
      "rawMarkdown": "Hi, thanks for the writeup.\n\nCan you please go into more detail on how you selected the features to apply the clustering algorithm on? I think this is a vital part that will be helpful to fully understand the data quality argument here.",
      "votes": null
    },
    {
      "id": "3266049",
      "postDate": "08/08/2025 09:56:02",
      "content": "<p>Thanks for sharing your brilliant solution — I found your data preparation approach particularly insightful.</p>\n<p>I have a question regarding the clustering and filtering process you described. When using hourly data (Y_M_D_H), it makes sense that each hour contains 60 rows (1-minute resolution), allowing you to cluster those rows and remove \"bad\" clusters based on validation performance.</p>\n<p>However, when switching to minute-level granularity (Y_M_D_H_M), each timestamp represents a single row. In this case, how did you apply clustering? Since each minute corresponds to only one data point, what was your strategy for clustering and filtering in that more granular setting?</p>\n<p>Would greatly appreciate any clarification!</p>",
      "rawMarkdown": "Thanks for sharing your brilliant solution — I found your data preparation approach particularly insightful.\n\nI have a question regarding the clustering and filtering process you described. When using hourly data (Y_M_D_H), it makes sense that each hour contains 60 rows (1-minute resolution), allowing you to cluster those rows and remove \"bad\" clusters based on validation performance.\n\nHowever, when switching to minute-level granularity (Y_M_D_H_M), each timestamp represents a single row. In this case, how did you apply clustering? Since each minute corresponds to only one data point, what was your strategy for clustering and filtering in that more granular setting?\n\nWould greatly appreciate any clarification!",
      "votes": null
    },
    {
      "id": "3266313",
      "postDate": "08/08/2025 19:18:17",
      "content": "<p>Wait, I thought for the hourly data, he clustered the hours according to some \"selected features\" and removed those clusters if it improved validation performance.<br>\nFrom what you are saying, you believe that he just clustered by hours and went through each hours one by one and dropped them if it improved validation performance. This doesn't seem to work in my experiments.</p>\n<p><a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a>, we will greatly appreciate if you could explain your approach in a bit more detail. This is probably the most fascinating approach in the whole competition and I am dying to learn more about this.</p>",
      "rawMarkdown": "Wait, I thought for the hourly data, he clustered the hours according to some \"selected features\" and removed those clusters if it improved validation performance.\nFrom what you are saying, you believe that he just clustered by hours and went through each hours one by one and dropped them if it improved validation performance. This doesn't seem to work in my experiments.\n\n@youneseloiarm, we will greatly appreciate if you could explain your approach in a bit more detail. This is probably the most fascinating approach in the whole competition and I am dying to learn more about this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3262911,
      "author_name": "ducnh279",
      "author_url": "",
      "post_date": "08/04/2025 13:14:15",
      "content": "<p>Hi, thanks for sharing your solution! </p>\n<p>It seems you did not perform cross-validation? if yes, what is your cross-validation strategy and its score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3262917,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/04/2025 13:24:04",
          "content": "<p>Linear models are among the oldest modeling techniques, commonly used in theoretical economic models and economic research in general. One key advantage is that they tend not to overfit as easily as many machine learning or deep learning models.</p>\n<p>However, you need to be careful with the features you feed into the model. If the features are meaningful and well-engineered, the linear model will likely perform well—it's that simple.</p>\n<p>As I mentioned earlier, the public leaderboard reflects 49% of the data, so I focused my optimization efforts accordingly.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3262921,
              "author_name": "ducnh279",
              "author_url": "",
              "post_date": "08/04/2025 13:32:26",
              "content": "<p>I’m a bit confused about your approach and your reply, but anyway - congrats on finishing 25th! 🎉</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3264118,
      "author_name": "anhquangphan",
      "author_url": "",
      "post_date": "08/06/2025 07:28:39",
      "content": "<p>Thank you for sharing your nice solution. May I ask how did you come up with the idea of removing certain group of row to make a better train set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3264311,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/06/2025 13:03:40",
          "content": "<p>Thanks for your kind words! Yes, the idea of removing certain groups of rows actually came from my experience working with financial time series data, where row-level filtering often helps improve generalization and stability. In those settings, some rows introduce noise or shift distribution in a way that harms model performance. So I applied the same intuition here. Curious—did you try something similar in your approach?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3264692,
              "author_name": "anhquangphan",
              "author_url": "",
              "post_date": "08/06/2025 20:19:17",
              "content": "<p>Wow that was interesting, I have only thought maybe removing outlier or sth simpler. For me I think having an ensemble of model where each model is trained on certain fixed sized window made my solution a little bit more robust. Other than that I mainly use XGBoost and MLP (which I found that less hidden neurons =&gt; more robust =&gt; better score in this scenario) on selected features. I certainly got a lot of luck when the private score was released. Anw, thank you for sharing!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3264849,
      "author_name": "archeron13",
      "author_url": "",
      "post_date": "08/06/2025 22:31:19",
      "content": "<p>Hi, thanks for the writeup.</p>\n<p>Can you please go into more detail on how you selected the features to apply the clustering algorithm on? I think this is a vital part that will be helpful to fully understand the data quality argument here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3266049,
      "author_name": "kingui",
      "author_url": "",
      "post_date": "08/08/2025 09:56:02",
      "content": "<p>Thanks for sharing your brilliant solution — I found your data preparation approach particularly insightful.</p>\n<p>I have a question regarding the clustering and filtering process you described. When using hourly data (Y_M_D_H), it makes sense that each hour contains 60 rows (1-minute resolution), allowing you to cluster those rows and remove \"bad\" clusters based on validation performance.</p>\n<p>However, when switching to minute-level granularity (Y_M_D_H_M), each timestamp represents a single row. In this case, how did you apply clustering? Since each minute corresponds to only one data point, what was your strategy for clustering and filtering in that more granular setting?</p>\n<p>Would greatly appreciate any clarification!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3266313,
          "author_name": "archeron13",
          "author_url": "",
          "post_date": "08/08/2025 19:18:17",
          "content": "<p>Wait, I thought for the hourly data, he clustered the hours according to some \"selected features\" and removed those clusters if it improved validation performance.<br>\nFrom what you are saying, you believe that he just clustered by hours and went through each hours one by one and dropped them if it improved validation performance. This doesn't seem to work in my experiments.</p>\n<p><a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a>, we will greatly appreciate if you could explain your approach in a bit more detail. This is probably the most fascinating approach in the whole competition and I am dying to learn more about this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3262903": "My approach focused more on **data preparation** than on overly complex modeling. I started by engineering a time-based feature, `Y_M_D_H`, extracted from the timestamp column, which had a 1-minute resolution.\n\nUsing this feature, I grouped the data into **hourly intervals** and applied **clustering techniques** on selected features to detect similarity patterns across rows.\n\nThen, through an **iterative filtering process**, I removed or retained specific rows based on their impact on the public leaderboard score. For instance, if removing a cluster of rows (belonging to a particular `Y_M_D_H` group) improved validation performance, I excluded them permanently. If performance dropped, I kept them. This process helped me build a **cleaned, optimized training set** for classical models.\n\nThe goal wasn't to create a perfect preprocessing pipeline, but to show that **clean, well-structured training data can outperform complex models trained on noisy or unfiltered data**.\n\nInterestingly:\n\n* For my **25th place solution**, I used a `1-hour` timestamp.\n* For my **2nd place private leaderboard solution**, I used a more granular `1-minute` timestamp, forming a `Y_M_D_H_M` feature.\n\nIn both cases, I used **the same model and features** — only the **data preparation strategy differed**.\n\n* 📄 [2nd Place Solution (Private LB)](https://www.kaggle.com/code/youneseloiarm/drw-2th-place-in-the-private-lb)\n* 📄 [25th Place Solution (Private LB)](https://www.kaggle.com/code/youneseloiarm/drw-25th-place-in-the-private-lb)",
    "3262911": "Hi, thanks for sharing your solution! \n\nIt seems you did not perform cross-validation? if yes, what is your cross-validation strategy and its score?",
    "3262917": "Linear models are among the oldest modeling techniques, commonly used in theoretical economic models and economic research in general. One key advantage is that they tend not to overfit as easily as many machine learning or deep learning models.\n\nHowever, you need to be careful with the features you feed into the model. If the features are meaningful and well-engineered, the linear model will likely perform well—it's that simple.\n\nAs I mentioned earlier, the public leaderboard reflects 49% of the data, so I focused my optimization efforts accordingly.",
    "3262921": "I’m a bit confused about your approach and your reply, but anyway - congrats on finishing 25th! 🎉",
    "3264118": "Thank you for sharing your nice solution. May I ask how did you come up with the idea of removing certain group of row to make a better train set.",
    "3264311": "Thanks for your kind words! Yes, the idea of removing certain groups of rows actually came from my experience working with financial time series data, where row-level filtering often helps improve generalization and stability. In those settings, some rows introduce noise or shift distribution in a way that harms model performance. So I applied the same intuition here. Curious—did you try something similar in your approach?",
    "3264692": "Wow that was interesting, I have only thought maybe removing outlier or sth simpler. For me I think having an ensemble of model where each model is trained on certain fixed sized window made my solution a little bit more robust. Other than that I mainly use XGBoost and MLP (which I found that less hidden neurons => more robust => better score in this scenario) on selected features. I certainly got a lot of luck when the private score was released. Anw, thank you for sharing!",
    "3264849": "Hi, thanks for the writeup.\n\nCan you please go into more detail on how you selected the features to apply the clustering algorithm on? I think this is a vital part that will be helpful to fully understand the data quality argument here.",
    "3266049": "Thanks for sharing your brilliant solution — I found your data preparation approach particularly insightful.\n\nI have a question regarding the clustering and filtering process you described. When using hourly data (Y_M_D_H), it makes sense that each hour contains 60 rows (1-minute resolution), allowing you to cluster those rows and remove \"bad\" clusters based on validation performance.\n\nHowever, when switching to minute-level granularity (Y_M_D_H_M), each timestamp represents a single row. In this case, how did you apply clustering? Since each minute corresponds to only one data point, what was your strategy for clustering and filtering in that more granular setting?\n\nWould greatly appreciate any clarification!",
    "3266313": "Wait, I thought for the hourly data, he clustered the hours according to some \"selected features\" and removed those clusters if it improved validation performance.\nFrom what you are saying, you believe that he just clustered by hours and went through each hours one by one and dropped them if it improved validation performance. This doesn't seem to work in my experiments.\n\n@youneseloiarm, we will greatly appreciate if you could explain your approach in a bit more detail. This is probably the most fascinating approach in the whole competition and I am dying to learn more about this."
  },
  "source": "meta"
}