{
  "id": 552516,
  "title": "Local Clustered Group Sampling - A Strategy for Small and Noisy Data",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552516",
  "author_name": "",
  "post_date": "2024-12-20T04:11:59.684165100Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all, thanks to the competition organizers and all participants, and I wish everyone a Merry Christmas &amp; Happy New Year.</p>\n<p>This competition was interesting, tackling a complex and composite classification problem with a relatively small dataset (~3000 samples). As discussed by many in the forums, the selection of the random seed seems to play a significant role when accuracy is low (even the final best predictions barely reached 0.5). This randomness was the issue I most wanted to address during the competition. While I don’t think I fully resolved it, I may have come a bit closer, so I’d like to share my thoughts here and hear what others think.</p>\n<p>I believe that random seeds in this competition effectively partitioned the test data into groups. In other words, if the training data happens to match the test data well, the scores are relatively higher; otherwise, they are lower. Thus, averaging results across multiple seeds can yield a stable (though not optimal) score. This hypothesis was confirmed by my experiments. My best score came from a single seed about a month ago, achieving 0.479 (PB: 0.46). However, using the same method with scores averaged across multiple seeds resulted in a slightly lower but more stable score. This stable score became my baseline for evaluating whether a model was effective.</p>\n<p>The concept is quite simple: when analyzing a test sample, instead of using all the training data, I use only the data points most similar to that test sample. In practice, I calculated the \"distance\" between data points using three metrics: 'manhattan', 'euclidean', and 'cosine'. For each metric, I tried to select an equal number of the closest points. I further filtered the data by excluding points within the same sii group that were too close to each other, due to the significant imbalance in the number of samples across sii groups. Ultimately, I ended up selecting about 100 data points. Using these selected points, I trained a custom classification model to predict the test sample. I named this method “Local Clustered Group Sampling,” suggested by ChatGPT.</p>\n<p>The core idea of this method is to transform the randomness of the random seed (used for training set sampling) into something that we can control—by identifying similar data points and building models based on them. After adopting this approach, the involvement of random seeds is effectively eliminated. Of course, there is still room for optimization in the sampling process.</p>\n<p>The drawbacks of this approach are obvious. Each test sample requires a custom-built model (in fact, two: one with LightGBM and one with XGBoost), making it highly time-consuming. Fine-tuning the model also became quite challenging. My final submission was sent just 24 hours before the deadline, and it took nearly 9 hours to get the results after submission.</p>\n<p>Based on my offline tests and the final results, I believe this method has some merit. Currently, it does not outperform a full-model approach (using all data) across the entire test set. However, for certain test samples (typically those with 50+ nearby points identified), it can achieve comparable scores to the full model. In the end, I combined the predictions of these points using a voting system between the two models, which stabilized my average score and improved it by about 0.005.</p>\n<p>In summary, this method has much room for improvement, and I am unsure if it can ultimately outperform a full-model approach. However, in scenarios with limited and noisy data, it might be a viable strategy worth exploring.</p>\n<p>For me, this is an interesting competition because it is very real. Most of the time, we are faced with insufficient and chaotic data, and hope to find some patterns~</p>",
  "messages": [
    {
      "id": "3076556",
      "postDate": "12/20/2024 04:11:59",
      "content": "<p>First of all, thanks to the competition organizers and all participants, and I wish everyone a Merry Christmas &amp; Happy New Year.</p>\n<p>This competition was interesting, tackling a complex and composite classification problem with a relatively small dataset (~3000 samples). As discussed by many in the forums, the selection of the random seed seems to play a significant role when accuracy is low (even the final best predictions barely reached 0.5). This randomness was the issue I most wanted to address during the competition. While I don’t think I fully resolved it, I may have come a bit closer, so I’d like to share my thoughts here and hear what others think.</p>\n<p>I believe that random seeds in this competition effectively partitioned the test data into groups. In other words, if the training data happens to match the test data well, the scores are relatively higher; otherwise, they are lower. Thus, averaging results across multiple seeds can yield a stable (though not optimal) score. This hypothesis was confirmed by my experiments. My best score came from a single seed about a month ago, achieving 0.479 (PB: 0.46). However, using the same method with scores averaged across multiple seeds resulted in a slightly lower but more stable score. This stable score became my baseline for evaluating whether a model was effective.</p>\n<p>The concept is quite simple: when analyzing a test sample, instead of using all the training data, I use only the data points most similar to that test sample. In practice, I calculated the \"distance\" between data points using three metrics: 'manhattan', 'euclidean', and 'cosine'. For each metric, I tried to select an equal number of the closest points. I further filtered the data by excluding points within the same sii group that were too close to each other, due to the significant imbalance in the number of samples across sii groups. Ultimately, I ended up selecting about 100 data points. Using these selected points, I trained a custom classification model to predict the test sample. I named this method “Local Clustered Group Sampling,” suggested by ChatGPT.</p>\n<p>The core idea of this method is to transform the randomness of the random seed (used for training set sampling) into something that we can control—by identifying similar data points and building models based on them. After adopting this approach, the involvement of random seeds is effectively eliminated. Of course, there is still room for optimization in the sampling process.</p>\n<p>The drawbacks of this approach are obvious. Each test sample requires a custom-built model (in fact, two: one with LightGBM and one with XGBoost), making it highly time-consuming. Fine-tuning the model also became quite challenging. My final submission was sent just 24 hours before the deadline, and it took nearly 9 hours to get the results after submission.</p>\n<p>Based on my offline tests and the final results, I believe this method has some merit. Currently, it does not outperform a full-model approach (using all data) across the entire test set. However, for certain test samples (typically those with 50+ nearby points identified), it can achieve comparable scores to the full model. In the end, I combined the predictions of these points using a voting system between the two models, which stabilized my average score and improved it by about 0.005.</p>\n<p>In summary, this method has much room for improvement, and I am unsure if it can ultimately outperform a full-model approach. However, in scenarios with limited and noisy data, it might be a viable strategy worth exploring.</p>\n<p>For me, this is an interesting competition because it is very real. Most of the time, we are faced with insufficient and chaotic data, and hope to find some patterns~</p>",
      "rawMarkdown": "First of all, thanks to the competition organizers and all participants, and I wish everyone a Merry Christmas & Happy New Year.\n\nThis competition was interesting, tackling a complex and composite classification problem with a relatively small dataset (~3000 samples). As discussed by many in the forums, the selection of the random seed seems to play a significant role when accuracy is low (even the final best predictions barely reached 0.5). This randomness was the issue I most wanted to address during the competition. While I don’t think I fully resolved it, I may have come a bit closer, so I’d like to share my thoughts here and hear what others think.\n\nI believe that random seeds in this competition effectively partitioned the test data into groups. In other words, if the training data happens to match the test data well, the scores are relatively higher; otherwise, they are lower. Thus, averaging results across multiple seeds can yield a stable (though not optimal) score. This hypothesis was confirmed by my experiments. My best score came from a single seed about a month ago, achieving 0.479 (PB: 0.46). However, using the same method with scores averaged across multiple seeds resulted in a slightly lower but more stable score. This stable score became my baseline for evaluating whether a model was effective.\n\nThe concept is quite simple: when analyzing a test sample, instead of using all the training data, I use only the data points most similar to that test sample. In practice, I calculated the \"distance\" between data points using three metrics: 'manhattan', 'euclidean', and 'cosine'. For each metric, I tried to select an equal number of the closest points. I further filtered the data by excluding points within the same sii group that were too close to each other, due to the significant imbalance in the number of samples across sii groups. Ultimately, I ended up selecting about 100 data points. Using these selected points, I trained a custom classification model to predict the test sample. I named this method “Local Clustered Group Sampling,” suggested by ChatGPT.\n\nThe core idea of this method is to transform the randomness of the random seed (used for training set sampling) into something that we can control—by identifying similar data points and building models based on them. After adopting this approach, the involvement of random seeds is effectively eliminated. Of course, there is still room for optimization in the sampling process.\n\nThe drawbacks of this approach are obvious. Each test sample requires a custom-built model (in fact, two: one with LightGBM and one with XGBoost), making it highly time-consuming. Fine-tuning the model also became quite challenging. My final submission was sent just 24 hours before the deadline, and it took nearly 9 hours to get the results after submission.\n\nBased on my offline tests and the final results, I believe this method has some merit. Currently, it does not outperform a full-model approach (using all data) across the entire test set. However, for certain test samples (typically those with 50+ nearby points identified), it can achieve comparable scores to the full model. In the end, I combined the predictions of these points using a voting system between the two models, which stabilized my average score and improved it by about 0.005.\n\nIn summary, this method has much room for improvement, and I am unsure if it can ultimately outperform a full-model approach. However, in scenarios with limited and noisy data, it might be a viable strategy worth exploring.\n\n\nFor me, this is an interesting competition because it is very real. Most of the time, we are faced with insufficient and chaotic data, and hope to find some patterns~",
      "votes": null
    },
    {
      "id": "3076581",
      "postDate": "12/20/2024 05:06:53",
      "content": "<p>This is a good idea to explore but isn't the idea of cherry-picking similar train-test samples creating selection bias?<br>\nAlso, if the dataset would be slightly bigger, say 5_000 rows (small otherwise but bigger than this one), then this will take too much time to complete <a href=\"https://www.kaggle.com/hy2017\" target=\"_blank\">@hy2017</a> <br>\nCongrats for the result!</p>",
      "rawMarkdown": "This is a good idea to explore but isn't the idea of cherry-picking similar train-test samples creating selection bias?\nAlso, if the dataset would be slightly bigger, say 5_000 rows (small otherwise but bigger than this one), then this will take too much time to complete @hy2017 \nCongrats for the result!",
      "votes": null
    },
    {
      "id": "3077019",
      "postDate": "12/20/2024 13:34:24",
      "content": "<p>I think \"selecting nearby points\" is the most challenging aspect of this method to fine-tune.<br>\nI basically follow consistent rules to select points. Under some rules, the selection can be highly divergent, resulting in very poor scores. However, when I achieve a more convergent selection result, not only does the score improve significantly, but an interesting phenomenon also emerges: it’s often possible to guess the final answer from the distribution of selected points.</p>\n<p>For example, if a certain test sample receives 125 points of predictions with distributions like [0, 1, 2, 3] corresponding to [93, 24, 8, 0], it’s almost certain that this sample will be classified as 0. However, there are also cases where the results appear ambiguous, such as [0, 105, 158, 12]. The correct answer for this group is 1, but the model consistently predicts 2.</p>\n<p>Therefore, i think the effectiveness of this method depends significantly on the quality of the rules used for point selection.</p>\n<p>Currently, my selection method involves gradually relaxing the proportion of features considered. For example:<br>\nIn the first stage, I select points using the top 50% of the most similar features.<br>\nIf there aren’t enough points, I proceed to the second stage, where I only consider the top 40%, and so on.</p>\n<p>As for the time issue, with approximately 2700 samples in the training dataset, each test point requires about 1.5 seconds of computation time on my computer.<br>\nSo, if there are 3000 test samples, I would need around 4500 seconds in total.<br>\nI believe there’s still significant room for optimization in this aspect.</p>",
      "rawMarkdown": "I think \"selecting nearby points\" is the most challenging aspect of this method to fine-tune.\nI basically follow consistent rules to select points. Under some rules, the selection can be highly divergent, resulting in very poor scores. However, when I achieve a more convergent selection result, not only does the score improve significantly, but an interesting phenomenon also emerges: it’s often possible to guess the final answer from the distribution of selected points.\n\nFor example, if a certain test sample receives 125 points of predictions with distributions like [0, 1, 2, 3] corresponding to [93, 24, 8, 0], it’s almost certain that this sample will be classified as 0. However, there are also cases where the results appear ambiguous, such as [0, 105, 158, 12]. The correct answer for this group is 1, but the model consistently predicts 2.\n\nTherefore, i think the effectiveness of this method depends significantly on the quality of the rules used for point selection.\n\nCurrently, my selection method involves gradually relaxing the proportion of features considered. For example:\nIn the first stage, I select points using the top 50% of the most similar features.\nIf there aren’t enough points, I proceed to the second stage, where I only consider the top 40%, and so on.\n\nAs for the time issue, with approximately 2700 samples in the training dataset, each test point requires about 1.5 seconds of computation time on my computer.\nSo, if there are 3000 test samples, I would need around 4500 seconds in total.\nI believe there’s still significant room for optimization in this aspect.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076581,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/20/2024 05:06:53",
      "content": "<p>This is a good idea to explore but isn't the idea of cherry-picking similar train-test samples creating selection bias?<br>\nAlso, if the dataset would be slightly bigger, say 5_000 rows (small otherwise but bigger than this one), then this will take too much time to complete <a href=\"https://www.kaggle.com/hy2017\" target=\"_blank\">@hy2017</a> <br>\nCongrats for the result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3077019,
      "author_name": "hy2017",
      "author_url": "",
      "post_date": "12/20/2024 13:34:24",
      "content": "<p>I think \"selecting nearby points\" is the most challenging aspect of this method to fine-tune.<br>\nI basically follow consistent rules to select points. Under some rules, the selection can be highly divergent, resulting in very poor scores. However, when I achieve a more convergent selection result, not only does the score improve significantly, but an interesting phenomenon also emerges: it’s often possible to guess the final answer from the distribution of selected points.</p>\n<p>For example, if a certain test sample receives 125 points of predictions with distributions like [0, 1, 2, 3] corresponding to [93, 24, 8, 0], it’s almost certain that this sample will be classified as 0. However, there are also cases where the results appear ambiguous, such as [0, 105, 158, 12]. The correct answer for this group is 1, but the model consistently predicts 2.</p>\n<p>Therefore, i think the effectiveness of this method depends significantly on the quality of the rules used for point selection.</p>\n<p>Currently, my selection method involves gradually relaxing the proportion of features considered. For example:<br>\nIn the first stage, I select points using the top 50% of the most similar features.<br>\nIf there aren’t enough points, I proceed to the second stage, where I only consider the top 40%, and so on.</p>\n<p>As for the time issue, with approximately 2700 samples in the training dataset, each test point requires about 1.5 seconds of computation time on my computer.<br>\nSo, if there are 3000 test samples, I would need around 4500 seconds in total.<br>\nI believe there’s still significant room for optimization in this aspect.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3076556": "First of all, thanks to the competition organizers and all participants, and I wish everyone a Merry Christmas & Happy New Year.\n\nThis competition was interesting, tackling a complex and composite classification problem with a relatively small dataset (~3000 samples). As discussed by many in the forums, the selection of the random seed seems to play a significant role when accuracy is low (even the final best predictions barely reached 0.5). This randomness was the issue I most wanted to address during the competition. While I don’t think I fully resolved it, I may have come a bit closer, so I’d like to share my thoughts here and hear what others think.\n\nI believe that random seeds in this competition effectively partitioned the test data into groups. In other words, if the training data happens to match the test data well, the scores are relatively higher; otherwise, they are lower. Thus, averaging results across multiple seeds can yield a stable (though not optimal) score. This hypothesis was confirmed by my experiments. My best score came from a single seed about a month ago, achieving 0.479 (PB: 0.46). However, using the same method with scores averaged across multiple seeds resulted in a slightly lower but more stable score. This stable score became my baseline for evaluating whether a model was effective.\n\nThe concept is quite simple: when analyzing a test sample, instead of using all the training data, I use only the data points most similar to that test sample. In practice, I calculated the \"distance\" between data points using three metrics: 'manhattan', 'euclidean', and 'cosine'. For each metric, I tried to select an equal number of the closest points. I further filtered the data by excluding points within the same sii group that were too close to each other, due to the significant imbalance in the number of samples across sii groups. Ultimately, I ended up selecting about 100 data points. Using these selected points, I trained a custom classification model to predict the test sample. I named this method “Local Clustered Group Sampling,” suggested by ChatGPT.\n\nThe core idea of this method is to transform the randomness of the random seed (used for training set sampling) into something that we can control—by identifying similar data points and building models based on them. After adopting this approach, the involvement of random seeds is effectively eliminated. Of course, there is still room for optimization in the sampling process.\n\nThe drawbacks of this approach are obvious. Each test sample requires a custom-built model (in fact, two: one with LightGBM and one with XGBoost), making it highly time-consuming. Fine-tuning the model also became quite challenging. My final submission was sent just 24 hours before the deadline, and it took nearly 9 hours to get the results after submission.\n\nBased on my offline tests and the final results, I believe this method has some merit. Currently, it does not outperform a full-model approach (using all data) across the entire test set. However, for certain test samples (typically those with 50+ nearby points identified), it can achieve comparable scores to the full model. In the end, I combined the predictions of these points using a voting system between the two models, which stabilized my average score and improved it by about 0.005.\n\nIn summary, this method has much room for improvement, and I am unsure if it can ultimately outperform a full-model approach. However, in scenarios with limited and noisy data, it might be a viable strategy worth exploring.\n\n\nFor me, this is an interesting competition because it is very real. Most of the time, we are faced with insufficient and chaotic data, and hope to find some patterns~",
    "3076581": "This is a good idea to explore but isn't the idea of cherry-picking similar train-test samples creating selection bias?\nAlso, if the dataset would be slightly bigger, say 5_000 rows (small otherwise but bigger than this one), then this will take too much time to complete @hy2017 \nCongrats for the result!",
    "3077019": "I think \"selecting nearby points\" is the most challenging aspect of this method to fine-tune.\nI basically follow consistent rules to select points. Under some rules, the selection can be highly divergent, resulting in very poor scores. However, when I achieve a more convergent selection result, not only does the score improve significantly, but an interesting phenomenon also emerges: it’s often possible to guess the final answer from the distribution of selected points.\n\nFor example, if a certain test sample receives 125 points of predictions with distributions like [0, 1, 2, 3] corresponding to [93, 24, 8, 0], it’s almost certain that this sample will be classified as 0. However, there are also cases where the results appear ambiguous, such as [0, 105, 158, 12]. The correct answer for this group is 1, but the model consistently predicts 2.\n\nTherefore, i think the effectiveness of this method depends significantly on the quality of the rules used for point selection.\n\nCurrently, my selection method involves gradually relaxing the proportion of features considered. For example:\nIn the first stage, I select points using the top 50% of the most similar features.\nIf there aren’t enough points, I proceed to the second stage, where I only consider the top 40%, and so on.\n\nAs for the time issue, with approximately 2700 samples in the training dataset, each test point requires about 1.5 seconds of computation time on my computer.\nSo, if there are 3000 test samples, I would need around 4500 seconds in total.\nI believe there’s still significant room for optimization in this aspect."
  },
  "source": "meta"
}