{
  "id": 498311,
  "title": "Is it worth giving this competition a try using just Kaggle resources?",
  "url": "/competitions/leash-BELKA/discussion/498311",
  "author_name": "Frenio Redeker",
  "post_date": "2024-04-27T18:38:29.782000",
  "votes": 15,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>I am thinking about joining this competition, but it looks like the training data is huge even after <a href=\"https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading\" target=\"_blank\">shrinking</a> the dataset significantly. Additionally, it seems like people who post public notebooks use only a fraction of the dataset for training and I am worried that Kaggle's limits on CPU RAM will make it impossible to compete. So I thought I'd just ask what people who are already competing think:</p>\n<p>Is it possible to compete in this competition if I am restricted to using Kaggle resources?</p>",
  "messages": [
    {
      "id": 2779627,
      "postDate": "2024-04-27T18:38:29.783Z",
      "content": "<p>Hi All,</p>\n<p>I am thinking about joining this competition, but it looks like the training data is huge even after <a href=\"https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading\" target=\"_blank\">shrinking</a> the dataset significantly. Additionally, it seems like people who post public notebooks use only a fraction of the dataset for training and I am worried that Kaggle's limits on CPU RAM will make it impossible to compete. So I thought I'd just ask what people who are already competing think:</p>\n<p>Is it possible to compete in this competition if I am restricted to using Kaggle resources?</p>",
      "rawMarkdown": "Hi All,\n\nI am thinking about joining this competition, but it looks like the training data is huge even after [shrinking](https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading) the dataset significantly. Additionally, it seems like people who post public notebooks use only a fraction of the dataset for training and I am worried that Kaggle's limits on CPU RAM will make it impossible to compete. So I thought I'd just ask what people who are already competing think:\n\nIs it possible to compete in this competition if I am restricted to using Kaggle resources?",
      "votes": 15
    },
    {
      "id": 2779715,
      "postDate": "2024-04-27T19:36:19.457Z",
      "content": "<p>It's possible and I'm doing it.</p>\n<p>It is a significant extra challenge, and requires a certain amount of extra work and sophistication. And handling frustration, lol. But even double or quadruple RAM on a local PC would still be a challenge, I think, so it's hard to say how big a difference it makes, it's just a lot more black and white on Kaggle notebooks. To an extent, everyone still needs to be careful about number of rows and number of features used for training.</p>\n<p>Kaggle tricks can include:</p>\n<ul>\n<li>Batch sizes.</li>\n<li>Search Kaggle notebooks for good tricks, since it's a common issue.</li>\n<li>Dask etc (but I don't have much of any experience yet)</li>\n<li>Batch sizes.</li>\n<li>Ensembles trained on less data (less rows and/or less features)</li>\n<li>Prep the data in one notebook and save it in the best format for the model (LightGBM.Dataset, for instance), then have a new notebook loads that binary and start training.</li>\n<li>Reduce data type size</li>\n<li>Batch sizes lol.</li>\n<li>Not entirely sure if/when it's necessary, or it might even depend on model, but I try to have all data at the same data size. Then it can be one np array. For example, if using fingerprint bit features, use ONLY bit features, (so adding other one-hot encoded features would work fine) don't combine it with any float32, or int16 features. Maybe instead train the fingerprint model alone, and ensemble the output, or even use the output as a feature in a different model.</li>\n</ul>",
      "rawMarkdown": "It's possible and I'm doing it.\n\nIt is a significant extra challenge, and requires a certain amount of extra work and sophistication. And handling frustration, lol. But even double or quadruple RAM on a local PC would still be a challenge, I think, so it's hard to say how big a difference it makes, it's just a lot more black and white on Kaggle notebooks. To an extent, everyone still needs to be careful about number of rows and number of features used for training.\n\nKaggle tricks can include:\n- Batch sizes.\n- Search Kaggle notebooks for good tricks, since it's a common issue.\n- Dask etc (but I don't have much of any experience yet)\n- Batch sizes.\n- Ensembles trained on less data (less rows and/or less features)\n- Prep the data in one notebook and save it in the best format for the model (LightGBM.Dataset, for instance), then have a new notebook loads that binary and start training.\n- Reduce data type size\n- Batch sizes lol.\n- Not entirely sure if/when it's necessary, or it might even depend on model, but I try to have all data at the same data size. Then it can be one np array. For example, if using fingerprint bit features, use ONLY bit features, (so adding other one-hot encoded features would work fine) don't combine it with any float32, or int16 features. Maybe instead train the fingerprint model alone, and ensemble the output, or even use the output as a feature in a different model.",
      "votes": 7,
      "replies": [
        {
          "id": 2779717,
          "postDate": "2024-04-27T19:37:28.190Z",
          "content": "<p>Haha, we posted seconds apart <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> . Thanks for confirming my hypothesis that with a well laid out plan you can potentially do very well in this competition using Kaggle compute. Well done getting top 5 LB score (04/27/2024) BTW! I'm curious to see your strategies at the end of this comp.</p>",
          "rawMarkdown": "Haha, we posted seconds apart @roberthatch . Thanks for confirming my hypothesis that with a well laid out plan you can potentially do very well in this competition using Kaggle compute. Well done getting top 5 LB score (04/27/2024) BTW! I'm curious to see your strategies at the end of this comp.",
          "votes": 4,
          "replies": [
            {
              "id": 2779731,
              "postDate": "2024-04-27T19:46:00.947Z",
              "content": "<p>Thanks for the shout-out. :)</p>\n<p>My first sub was 6th place at 0.601! (04/22/2024) I was more than a little surprised, lol. I think the current \"secret sauce\" will make for an interesting read, if nothing else. :)</p>\n<p>Even all the people using local compute are still complaining about data size and training time and inference time, so maybe in some ways I had an advantage in planning for it from the start…?</p>",
              "rawMarkdown": "Thanks for the shout-out. :)\n\nMy first sub was 6th place at 0.601! (04/22/2024) I was more than a little surprised, lol. I think the current \"secret sauce\" will make for an interesting read, if nothing else. :)\n\nEven all the people using local compute are still complaining about data size and training time and inference time, so maybe in some ways I had an advantage in planning for it from the start...?",
              "votes": 4
            },
            {
              "id": 2779740,
              "postDate": "2024-04-27T19:54:51.363Z",
              "content": "<p>Yeah, I noticed you made quick work to hop into 6th! Nicely done! My models up to now have been exploring different fingerprints and train/test sizes with fairly naive scaffold splits. Doing a bit more rational stuff now, so hopefully I can find some secret sauce too :).</p>\n<p>Best of luck!</p>",
              "rawMarkdown": "Yeah, I noticed you made quick work to hop into 6th! Nicely done! My models up to now have been exploring different fingerprints and train/test sizes with fairly naive scaffold splits. Doing a bit more rational stuff now, so hopefully I can find some secret sauce too :).\n\nBest of luck!",
              "votes": 2
            }
          ]
        },
        {
          "id": 2779743,
          "postDate": "2024-04-27T19:55:30.850Z",
          "content": "<p>Congrats on doing so well so far! This is very encouraging and thank you for sharing some of the tricks you're using or might be using in the future. You definitely got me closer to deciding in favor of joining the competition.</p>",
          "rawMarkdown": "Congrats on doing so well so far! This is very encouraging and thank you for sharing some of the tricks you're using or might be using in the future. You definitely got me closer to deciding in favor of joining the competition.",
          "votes": 2
        },
        {
          "id": 2779834,
          "postDate": "2024-04-27T20:56:33.223Z",
          "content": "<p>Just to add to your list, some data features like \"ecfp\" can be stored as sparse matrix.</p>",
          "rawMarkdown": "Just to add to your list, some data features like \"ecfp\" can be stored as sparse matrix.",
          "votes": 3,
          "replies": [
            {
              "id": 2780066,
              "postDate": "2024-04-28T02:54:29.563Z",
              "content": "<p>Thanks I'll have to experiment with that</p>",
              "rawMarkdown": "Thanks I'll have to experiment with that",
              "votes": 3
            },
            {
              "id": 2781553,
              "postDate": "2024-04-28T20:42:40.117Z",
              "content": "<p>if you are using sparse matrix, do remember to check for bug<br>\n<a href=\"https://github.com/dmlc/xgboost/issues/7729\" target=\"_blank\">https://github.com/dmlc/xgboost/issues/7729</a></p>",
              "rawMarkdown": "if you are using sparse matrix, do remember to check for bug\nhttps://github.com/dmlc/xgboost/issues/7729\n",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2779969,
      "postDate": "2024-04-28T00:41:05.340Z",
      "content": "<p>the more difficult it is, the more you learn.</p>\n<p>or if you believe in Multi-armed bandit model:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa9d2c9aa4b92c4757c88549ac3de098a%2Froad.not_.taken_.jpg?generation=1714264922100297&amp;alt=media\"></p>",
      "rawMarkdown": "the more difficult it is, the more you learn.\n\nor if you believe in Multi-armed bandit model:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa9d2c9aa4b92c4757c88549ac3de098a%2Froad.not_.taken_.jpg?generation=1714264922100297&alt=media)",
      "votes": 3
    },
    {
      "id": 2779714,
      "postDate": "2024-04-27T19:35:57.087Z",
      "content": "<p>I think it depends on what you mean by \"try\". If you want to come and learn some cheminformatics, machine learning, and data wrangling I think Kaggle resources are more than enough. However, it's true that having more compute power will help you be more competitive in this competition. That's not to say it's impossible to do well with just Kaggle compute and I suspect with a good sampling strategy and well laid out experimental plan, you could have a good score in this competition without using any external computing power.</p>",
      "rawMarkdown": "I think it depends on what you mean by \"try\". If you want to come and learn some cheminformatics, machine learning, and data wrangling I think Kaggle resources are more than enough. However, it's true that having more compute power will help you be more competitive in this competition. That's not to say it's impossible to do well with just Kaggle compute and I suspect with a good sampling strategy and well laid out experimental plan, you could have a good score in this competition without using any external computing power.",
      "votes": 3,
      "replies": [
        {
          "id": 2779720,
          "postDate": "2024-04-27T19:39:46.843Z",
          "content": "<blockquote>\n  <p>you could have a good score in this competition without using any external computing power.</p>\n</blockquote>\n<p>I've proven it's possible so far… :) But end of competition is a very different animal than early public LB score.</p>",
          "rawMarkdown": "> you could have a good score in this competition without using any external computing power.\n\nI've proven it's possible so far... :) But end of competition is a very different animal than early public LB score.",
          "votes": 3,
          "replies": [
            {
              "id": 2779722,
              "postDate": "2024-04-27T19:42:08.940Z",
              "content": "<p>Completely true we'll have to see how the competition evolves, but it's encouraging and tbh I find it exciting :)</p>",
              "rawMarkdown": "Completely true we'll have to see how the competition evolves, but it's encouraging and tbh I find it exciting :)",
              "votes": 3
            },
            {
              "id": 2780254,
              "postDate": "2024-04-28T05:30:08.530Z",
              "content": "<p>i believe the end score will be around 0.70 to 0.75 at the end of competition.<br>\ni recall the last enzyme competition, some kaggle are uses MD (molecular simulation).<br>\nHope to see different methods at the end …</p>",
              "rawMarkdown": "i believe the end score will be around 0.70 to 0.75 at the end of competition.\ni recall the last enzyme competition, some kaggle are uses MD (molecular simulation).\nHope to see different methods at the end ...",
              "votes": 3
            },
            {
              "id": 2781531,
              "postDate": "2024-04-28T20:29:19.037Z",
              "content": "<p>That sounds reasonable to me. I'm really hoping that we all do really well with the triazine molecules in the test set, but really learn from how we do against the non-triazines. I'm hopeful that this portion will have varied approaches. There's a lot to try and many ways to imagine setting a validation set for this.</p>",
              "rawMarkdown": "That sounds reasonable to me. I'm really hoping that we all do really well with the triazine molecules in the test set, but really learn from how we do against the non-triazines. I'm hopeful that this portion will have varied approaches. There's a lot to try and many ways to imagine setting a validation set for this.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2779760,
          "postDate": "2024-04-27T20:01:01.087Z",
          "content": "<p>Thank you! I'm sure there is a lot to learn in any case and looks like <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> has already proven your point about the possibility to be competitive with a good strategy and panning.</p>",
          "rawMarkdown": "Thank you! I'm sure there is a lot to learn in any case and looks like @roberthatch has already proven your point about the possibility to be competitive with a good strategy and panning.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2779880,
      "postDate": "2024-04-27T21:36:51.720Z",
      "content": "<p>Just use TPU</p>",
      "rawMarkdown": "Just use TPU",
      "votes": 2,
      "replies": [
        {
          "id": 2780002,
          "postDate": "2024-04-28T01:12:05.190Z",
          "content": "<p>I haven't worked with TPU so far, so I am not sure how that will solve the memory issues that I am worried about. Thank you for the hint, though! I will look into using TPU if I choose to join the competition.</p>",
          "rawMarkdown": "I haven't worked with TPU so far, so I am not sure how that will solve the memory issues that I am worried about. Thank you for the hint, though! I will look into using TPU if I choose to join the competition.",
          "replies": [
            {
              "id": 2780561,
              "postDate": "2024-04-28T09:53:08.537Z",
              "content": "<p>'How that will solve the memory issues'<br>\nOver 300GB ram, 128GB VRAM, over 90 CPU cores…</p>",
              "rawMarkdown": "'How that will solve the memory issues'\nOver 300GB ram, 128GB VRAM, over 90 CPU cores...",
              "votes": 1
            },
            {
              "id": 2781255,
              "postDate": "2024-04-28T17:10:24.273Z",
              "content": "<p>Wow, I had no idea! Thanks for the clarification.</p>",
              "rawMarkdown": "Wow, I had no idea! Thanks for the clarification."
            }
          ]
        }
      ]
    },
    {
      "id": 2779827,
      "postDate": "2024-04-27T20:50:27.310Z",
      "content": "<p>Since it is not a code competition and with that amount of data… competing is going to be complicated. A strong background in medicinal chemistry might also be helpful. But kaggle is not just about competing.</p>",
      "rawMarkdown": "Since it is not a code competition and with that amount of data... competing is going to be complicated. A strong background in medicinal chemistry might also be helpful. But kaggle is not just about competing.",
      "votes": 2,
      "replies": [
        {
          "id": 2779996,
          "postDate": "2024-04-28T01:07:54Z",
          "content": "<p>Yes, thanks for the reminder! I am mostly here to learn, but competitiveness is a great motivator for me.</p>",
          "rawMarkdown": "Yes, thanks for the reminder! I am mostly here to learn, but competitiveness is a great motivator for me.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2784600,
      "postDate": "2024-04-30T11:40:28.483Z",
      "content": "<p>That is same question I asked in my post. can just Kaggle resources like CPU, GPU and RAM would be sufficient or one have to write code on his/ her local workstation with sufficient computing resources and then upload model or its predictions and submit it. <br>\nCan anyone share links related to this. </p>",
      "rawMarkdown": "That is same question I asked in my post. can just Kaggle resources like CPU, GPU and RAM would be sufficient or one have to write code on his/ her local workstation with sufficient computing resources and then upload model or its predictions and submit it. \nCan anyone share links related to this. "
    },
    {
      "id": 2779685,
      "postDate": "2024-04-27T19:04:27.107Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2779771,
          "postDate": "2024-04-27T20:05:55.003Z",
          "content": "<p>Thank you for giving me your honest opinion, since that's what I was hoping to get when I opened this topic! I am sure there is a lot to learn in the competitions you suggest, but at the moment I am most interested in working on chemistry and biology related problems.</p>",
          "rawMarkdown": "Thank you for giving me your honest opinion, since that's what I was hoping to get when I opened this topic! I am sure there is a lot to learn in the competitions you suggest, but at the moment I am most interested in working on chemistry and biology related problems.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2779715,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2024-04-27T19:36:19.457000",
      "content": "<p>It's possible and I'm doing it.</p>\n<p>It is a significant extra challenge, and requires a certain amount of extra work and sophistication. And handling frustration, lol. But even double or quadruple RAM on a local PC would still be a challenge, I think, so it's hard to say how big a difference it makes, it's just a lot more black and white on Kaggle notebooks. To an extent, everyone still needs to be careful about number of rows and number of features used for training.</p>\n<p>Kaggle tricks can include:</p>\n<ul>\n<li>Batch sizes.</li>\n<li>Search Kaggle notebooks for good tricks, since it's a common issue.</li>\n<li>Dask etc (but I don't have much of any experience yet)</li>\n<li>Batch sizes.</li>\n<li>Ensembles trained on less data (less rows and/or less features)</li>\n<li>Prep the data in one notebook and save it in the best format for the model (LightGBM.Dataset, for instance), then have a new notebook loads that binary and start training.</li>\n<li>Reduce data type size</li>\n<li>Batch sizes lol.</li>\n<li>Not entirely sure if/when it's necessary, or it might even depend on model, but I try to have all data at the same data size. Then it can be one np array. For example, if using fingerprint bit features, use ONLY bit features, (so adding other one-hot encoded features would work fine) don't combine it with any float32, or int16 features. Maybe instead train the fingerprint model alone, and ensemble the output, or even use the output as a feature in a different model.</li>\n</ul>",
      "votes": 7,
      "replies": [
        {
          "id": 2779717,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-04-27T19:37:28.190000",
          "content": "<p>Haha, we posted seconds apart <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> . Thanks for confirming my hypothesis that with a well laid out plan you can potentially do very well in this competition using Kaggle compute. Well done getting top 5 LB score (04/27/2024) BTW! I'm curious to see your strategies at the end of this comp.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2779731,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-04-27T19:46:00.947000",
              "content": "<p>Thanks for the shout-out. :)</p>\n<p>My first sub was 6th place at 0.601! (04/22/2024) I was more than a little surprised, lol. I think the current \"secret sauce\" will make for an interesting read, if nothing else. :)</p>\n<p>Even all the people using local compute are still complaining about data size and training time and inference time, so maybe in some ways I had an advantage in planning for it from the start…?</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2779740,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-04-27T19:54:51.363000",
              "content": "<p>Yeah, I noticed you made quick work to hop into 6th! Nicely done! My models up to now have been exploring different fingerprints and train/test sizes with fairly naive scaffold splits. Doing a bit more rational stuff now, so hopefully I can find some secret sauce too :).</p>\n<p>Best of luck!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2779743,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2024-04-27T19:55:30.850000",
          "content": "<p>Congrats on doing so well so far! This is very encouraging and thank you for sharing some of the tricks you're using or might be using in the future. You definitely got me closer to deciding in favor of joining the competition.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2779834,
          "author_name": "Ricardo Colomer",
          "author_url": "",
          "post_date": "2024-04-27T20:56:33.223000",
          "content": "<p>Just to add to your list, some data features like \"ecfp\" can be stored as sparse matrix.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2780066,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2024-04-28T02:54:29.563000",
              "content": "<p>Thanks I'll have to experiment with that</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2781553,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-28T20:42:40.117000",
              "content": "<p>if you are using sparse matrix, do remember to check for bug<br>\n<a href=\"https://github.com/dmlc/xgboost/issues/7729\" target=\"_blank\">https://github.com/dmlc/xgboost/issues/7729</a></p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2779969,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-28T00:41:05.340000",
      "content": "<p>the more difficult it is, the more you learn.</p>\n<p>or if you believe in Multi-armed bandit model:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa9d2c9aa4b92c4757c88549ac3de098a%2Froad.not_.taken_.jpg?generation=1714264922100297&amp;alt=media\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2779714,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "2024-04-27T19:35:57.087000",
      "content": "<p>I think it depends on what you mean by \"try\". If you want to come and learn some cheminformatics, machine learning, and data wrangling I think Kaggle resources are more than enough. However, it's true that having more compute power will help you be more competitive in this competition. That's not to say it's impossible to do well with just Kaggle compute and I suspect with a good sampling strategy and well laid out experimental plan, you could have a good score in this competition without using any external computing power.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2779720,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2024-04-27T19:39:46.843000",
          "content": "<blockquote>\n  <p>you could have a good score in this competition without using any external computing power.</p>\n</blockquote>\n<p>I've proven it's possible so far… :) But end of competition is a very different animal than early public LB score.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2779722,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-04-27T19:42:08.940000",
              "content": "<p>Completely true we'll have to see how the competition evolves, but it's encouraging and tbh I find it exciting :)</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2780254,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-28T05:30:08.530000",
              "content": "<p>i believe the end score will be around 0.70 to 0.75 at the end of competition.<br>\ni recall the last enzyme competition, some kaggle are uses MD (molecular simulation).<br>\nHope to see different methods at the end …</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2781531,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-04-28T20:29:19.037000",
              "content": "<p>That sounds reasonable to me. I'm really hoping that we all do really well with the triazine molecules in the test set, but really learn from how we do against the non-triazines. I'm hopeful that this portion will have varied approaches. There's a lot to try and many ways to imagine setting a validation set for this.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2779760,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2024-04-27T20:01:01.087000",
          "content": "<p>Thank you! I'm sure there is a lot to learn in any case and looks like <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> has already proven your point about the possibility to be competitive with a good strategy and panning.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2779880,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-04-27T21:36:51.720000",
      "content": "<p>Just use TPU</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2780002,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2024-04-28T01:12:05.190000",
          "content": "<p>I haven't worked with TPU so far, so I am not sure how that will solve the memory issues that I am worried about. Thank you for the hint, though! I will look into using TPU if I choose to join the competition.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2780561,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-04-28T09:53:08.537000",
              "content": "<p>'How that will solve the memory issues'<br>\nOver 300GB ram, 128GB VRAM, over 90 CPU cores…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2781255,
              "author_name": "Frenio Redeker",
              "author_url": "",
              "post_date": "2024-04-28T17:10:24.273000",
              "content": "<p>Wow, I had no idea! Thanks for the clarification.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2779827,
      "author_name": "Ricardo Colomer",
      "author_url": "",
      "post_date": "2024-04-27T20:50:27.310000",
      "content": "<p>Since it is not a code competition and with that amount of data… competing is going to be complicated. A strong background in medicinal chemistry might also be helpful. But kaggle is not just about competing.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2779996,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2024-04-28T01:07:54",
          "content": "<p>Yes, thanks for the reminder! I am mostly here to learn, but competitiveness is a great motivator for me.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2784600,
      "author_name": "MTP",
      "author_url": "",
      "post_date": "2024-04-30T11:40:28.483000",
      "content": "<p>That is same question I asked in my post. can just Kaggle resources like CPU, GPU and RAM would be sufficient or one have to write code on his/ her local workstation with sufficient computing resources and then upload model or its predictions and submit it. <br>\nCan anyone share links related to this. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2779685,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-27T19:04:27.107000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 2779771,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2024-04-27T20:05:55.003000",
          "content": "<p>Thank you for giving me your honest opinion, since that's what I was hoping to get when I opened this topic! I am sure there is a lot to learn in the competitions you suggest, but at the moment I am most interested in working on chemistry and biology related problems.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2779627": "Hi All,\n\nI am thinking about joining this competition, but it looks like the training data is huge even after [shrinking](https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading) the dataset significantly. Additionally, it seems like people who post public notebooks use only a fraction of the dataset for training and I am worried that Kaggle's limits on CPU RAM will make it impossible to compete. So I thought I'd just ask what people who are already competing think:\n\nIs it possible to compete in this competition if I am restricted to using Kaggle resources?",
    "2779715": "It's possible and I'm doing it.\n\nIt is a significant extra challenge, and requires a certain amount of extra work and sophistication. And handling frustration, lol. But even double or quadruple RAM on a local PC would still be a challenge, I think, so it's hard to say how big a difference it makes, it's just a lot more black and white on Kaggle notebooks. To an extent, everyone still needs to be careful about number of rows and number of features used for training.\n\nKaggle tricks can include:\n- Batch sizes.\n- Search Kaggle notebooks for good tricks, since it's a common issue.\n- Dask etc (but I don't have much of any experience yet)\n- Batch sizes.\n- Ensembles trained on less data (less rows and/or less features)\n- Prep the data in one notebook and save it in the best format for the model (LightGBM.Dataset, for instance), then have a new notebook loads that binary and start training.\n- Reduce data type size\n- Batch sizes lol.\n- Not entirely sure if/when it's necessary, or it might even depend on model, but I try to have all data at the same data size. Then it can be one np array. For example, if using fingerprint bit features, use ONLY bit features, (so adding other one-hot encoded features would work fine) don't combine it with any float32, or int16 features. Maybe instead train the fingerprint model alone, and ensemble the output, or even use the output as a feature in a different model.",
    "2779969": "the more difficult it is, the more you learn.\n\nor if you believe in Multi-armed bandit model:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa9d2c9aa4b92c4757c88549ac3de098a%2Froad.not_.taken_.jpg?generation=1714264922100297&alt=media)",
    "2779714": "I think it depends on what you mean by \"try\". If you want to come and learn some cheminformatics, machine learning, and data wrangling I think Kaggle resources are more than enough. However, it's true that having more compute power will help you be more competitive in this competition. That's not to say it's impossible to do well with just Kaggle compute and I suspect with a good sampling strategy and well laid out experimental plan, you could have a good score in this competition without using any external computing power.",
    "2779880": "Just use TPU",
    "2779827": "Since it is not a code competition and with that amount of data... competing is going to be complicated. A strong background in medicinal chemistry might also be helpful. But kaggle is not just about competing.",
    "2784600": "That is same question I asked in my post. can just Kaggle resources like CPU, GPU and RAM would be sufficient or one have to write code on his/ her local workstation with sufficient computing resources and then upload model or its predictions and submit it. \nCan anyone share links related to this. ",
    "2779685": ""
  }
}