{
  "id": 549329,
  "title": "Full Training on Kaggle possible for the GPU Poor?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/549329",
  "author_name": "Seqaeon",
  "post_date": "2024-12-01T20:29:31.634000",
  "votes": 1,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Is there anyone here that has figured out a way to train all or most of the data on kaggle for Neural Networks? If so would you be willing to share your method?</p>",
  "messages": [
    {
      "id": 3060858,
      "postDate": "2024-12-02T07:09:34.933Z",
      "content": "<p>its possible, i train all my model on data from partition 4 to 9, and it hardly takes 16 Gib of ram. use GPU, and load data with polars.</p>\n<p>In my tests pandas is not able to handle the data efficiently as polars is able to, and dont do splits with sklearn functions, do it manually, with numpy.</p>",
      "rawMarkdown": "its possible, i train all my model on data from partition 4 to 9, and it hardly takes 16 Gib of ram. use GPU, and load data with polars.\n\nIn my tests pandas is not able to handle the data efficiently as polars is able to, and dont do splits with sklearn functions, do it manually, with numpy.\n \n",
      "votes": 1,
      "replies": [
        {
          "id": 3060868,
          "postDate": "2024-12-02T07:21:28.283Z",
          "content": "<p>I was talking more like 8-10 out of the partition. But also I still had trouble running 6 partitions through an MLP, 12 hours and it still didn't finish training 100 epochs, only trained  11 epochs. I load the data with polars but use pandas for some processing operations.</p>\n<p>Anyways, could you please share your template for it. You can remove the model or replace it with a simple MLP one, just need to see how you handle the data as it goes through the model. </p>",
          "rawMarkdown": "I was talking more like 8-10 out of the partition. But also I still had trouble running 6 partitions through an MLP, 12 hours and it still didn't finish training 100 epochs, only trained  11 epochs. I load the data with polars but use pandas for some processing operations.\n\nAnyways, could you please share your template for it. You can remove the model or replace it with a simple MLP one, just need to see how you handle the data as it goes through the model. ",
          "replies": [
            {
              "id": 3060877,
              "postDate": "2024-12-02T07:44:51.873Z",
              "content": "<p>yeah training time is a problem on single or even on double t4 gpu's</p>\n<p>to look cool i used the emojis !!</p>\n<pre><code> torch\n torch.nn  nn\n torch.optim  optim\n torch.nn.functional  F\n math\n time\n numpy  np\n matplotlib.pyplot  plt\n\n\n\n ():\n    ss_total = ((y_true - y_true.mean())**).()\n    ss_residual = ((y_true - y_pred)**).()\n      - ss_residual / ss_total\n torch.utils.data  DataLoader, TensorDataset, Subset\n\n ():\n    \n    ()\n    indices = np.arange(X.shape[])  \n    \n\n    \n    train_size = ( * (indices))  \n    val_size = (indices) - train_size  \n\n    train_indices = indices[:train_size]\n    val_indices = indices[train_size:]\n\n    \n    ()\n    X_tensor = torch.FloatTensor(X).to(device)\n    y_tensor = torch.FloatTensor(y).to(device)\n     X,y\n    dataset = TensorDataset(X_tensor, y_tensor)\n    ()\n    \n    train_dataset = Subset(dataset, train_indices)\n    val_dataset = Subset(dataset, val_indices)\n\n    \n    train_loader = DataLoader(train_dataset, batch_size=batch_size, shuffle=)\n    val_loader = DataLoader(val_dataset, batch_size=batch_size, shuffle=)\n\n    \n    ()\n    model= NMDL(,, ) \n\n    \n    criterion = nn.MSELoss()\n    optimizer = optim.Adam(model.parameters(), lr=)\n    \n    best_val_loss = ()\n    patience_counter = \n    best_model_state = \n\n    \n    ()\n    start_time = time.time()\n\n    train_losses, val_losses = [], []\n    train_r2_scores, val_r2_scores = [], []\n\n     epoch  (num_epochs):\n        \n        ()\n        model.train()\n        epoch_train_losses = []\n        epoch_train_preds = []\n        epoch_train_targets = []\n\n         batch_X, batch_y  train_loader:\n            optimizer.zero_grad()\n\n            \n            train_pred = model(batch_X).squeeze()\n            loss = criterion(train_pred, batch_y)\n\n            \n            loss.backward()\n            optimizer.step()\n\n            \n            epoch_train_losses.append(loss.item())\n            epoch_train_preds.append(train_pred.detach())\n            epoch_train_targets.append(batch_y)\n\n        \n        train_loss = np.mean(epoch_train_losses)\n        train_preds = torch.cat(epoch_train_preds)\n        train_targets = torch.cat(epoch_train_targets)\n        train_r2 = custom_r2_score(train_targets, train_preds)\n\n        \n        ()\n        model.()\n        epoch_val_losses = []\n        epoch_val_preds = []\n        epoch_val_targets = []\n\n         torch.no_grad():\n             batch_X, batch_y  val_loader:\n                val_pred = model(batch_X).squeeze()\n                val_loss = criterion(val_pred, batch_y)\n\n                \n                epoch_val_losses.append(val_loss.item())\n                epoch_val_preds.append(val_pred)\n                epoch_val_targets.append(batch_y)\n\n        \n        val_loss = np.mean(epoch_val_losses)\n        val_preds = torch.cat(epoch_val_preds)\n        val_targets = torch.cat(epoch_val_targets)\n        val_r2 = custom_r2_score(val_targets, val_preds)\n\n        \n        (\n              \n              )\n\n        \n         val_loss &lt; best_val_loss:\n            best_val_loss = val_loss\n            patience_counter = \n            best_model_state = model.state_dict()\n            torch.save(best_model_state, )\n            ()\n        :\n            patience_counter += \n\n         patience_counter &gt;= patience:\n            ()\n            \n\n        \n        train_losses.append(train_loss)\n        val_losses.append(val_loss)\n        train_r2_scores.append(train_r2)\n        val_r2_scores.append(val_r2)\n\n    \n    model.load_state_dict(best_model_state)\n\n    \n    total_time = time.time() - start_time\n    ()\n\n    \n    plt.figure(figsize=(, ))\n    plt.subplot(, , )\n    plt.plot(train_losses, label=)\n    plt.plot(val_losses, label=)\n    plt.legend()\n    plt.title()\n\n    plt.subplot(, , )\n    plt.plot(train_r2_scores, label=)\n    plt.plot(val_r2_scores, label=)\n    plt.legend()\n    plt.title()\n\n    plt.show()\n\n\n __name__ == :\n    \n    batch_size = \n    seq_len = \n    input_dim =   \n    ()\n    df = pl.read_parquet().(pl.col()&gt;=).fill_null()\n\n    \n    X = df[[  i  ()]+[,]].to_numpy()\n    y = df[].to_numpy()\n    ()\n     df\n    device = torch.device(  torch.cuda.is_available()  )\n\n    \n    train_hybrid_model(X, y, device)\n</code></pre>",
              "rawMarkdown": "yeah training time is a problem on single or even on double t4 gpu's\n \nto look cool i used the emojis !!\n\n```\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\nimport torch.nn.functional as F\nimport math\nimport time\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n\n\ndef custom_r2_score(y_true, y_pred):\n    ss_total = ((y_true - y_true.mean())**2).sum()\n    ss_residual = ((y_true - y_pred)**2).sum()\n    return 1 - ss_residual / ss_total\nfrom torch.utils.data import DataLoader, TensorDataset, Subset\n# Training function with early stopping and verbose output\ndef train_hybrid_model(X, y, device, patience=5, num_epochs=50, batch_size=1024):\n    # Data split\n    print('🔄 Splitting data into training and validation sets...')\n    indices = np.arange(X.shape[0])  # Create an array of indices\n    # np.random.shuffle(indices)  \n    \n    # 2. Split into 80% train, 20% validation\n    train_size = int(0.8 * len(indices))  # 80% for training\n    val_size = len(indices) - train_size  # 20% for validation\n    \n    train_indices = indices[:train_size]\n    val_indices = indices[train_size:]\n    \n    # 3. Create TensorDataset for the entire dataset (Lazy Loading will be used in Subset)\n    print('🔧 Converting data to tensors (with lazy loading)...')\n    X_tensor = torch.FloatTensor(X).to(device)\n    y_tensor = torch.FloatTensor(y).to(device)\n    del X,y\n    dataset = TensorDataset(X_tensor, y_tensor)\n    print('📦 Creating data loaders...')\n    # 4. Create Subset for training and validation using the indices (no extra memory copy)\n    train_dataset = Subset(dataset, train_indices)\n    val_dataset = Subset(dataset, val_indices)\n    \n    # 5. Create DataLoader with batch processing (only loads data when needed)\n    train_loader = DataLoader(train_dataset, batch_size=batch_size, shuffle=True)\n    val_loader = DataLoader(val_dataset, batch_size=batch_size, shuffle=False)\n    \n    #\n    print('🧠 Initializing model...')\n    model= NMDL(100,96, 2040) #this is the model loading\n    \n    # Loss and optimizer\n    criterion = nn.MSELoss()\n    optimizer = optim.Adam(model.parameters(), lr=0.001)\n    # Early Stopping and best model saving\n    best_val_loss = float('inf')\n    patience_counter = 0\n    best_model_state = None\n    \n    # Training loop\n    print(\"\\n🏋️ Starting Training with Early Stopping...\")\n    start_time = time.time()\n    \n    train_losses, val_losses = [], []\n    train_r2_scores, val_r2_scores = [], []\n    \n    for epoch in range(num_epochs):\n        # Training phase\n        print(f\"🔄 Epoch {epoch+1}/{num_epochs}: Starting training phase...\")\n        model.train()\n        epoch_train_losses = []\n        epoch_train_preds = []\n        epoch_train_targets = []\n        \n        for batch_X, batch_y in train_loader:\n            optimizer.zero_grad()\n            \n            # Forward pass\n            train_pred = model(batch_X).squeeze()\n            loss = criterion(train_pred, batch_y)\n            \n            # Backward pass\n            loss.backward()\n            optimizer.step()\n            \n            # Collect batch results\n            epoch_train_losses.append(loss.item())\n            epoch_train_preds.append(train_pred.detach())\n            epoch_train_targets.append(batch_y)\n        \n        # Aggregate training results\n        train_loss = np.mean(epoch_train_losses)\n        train_preds = torch.cat(epoch_train_preds)\n        train_targets = torch.cat(epoch_train_targets)\n        train_r2 = custom_r2_score(train_targets, train_preds)\n        \n        # Validation phase\n        print(f\"✅ Epoch {epoch+1}/{num_epochs}: Starting validation phase...\")\n        model.eval()\n        epoch_val_losses = []\n        epoch_val_preds = []\n        epoch_val_targets = []\n        \n        with torch.no_grad():\n            for batch_X, batch_y in val_loader:\n                val_pred = model(batch_X).squeeze()\n                val_loss = criterion(val_pred, batch_y)\n                \n                # Collect batch results\n                epoch_val_losses.append(val_loss.item())\n                epoch_val_preds.append(val_pred)\n                epoch_val_targets.append(batch_y)\n        \n        # Aggregate validation results\n        val_loss = np.mean(epoch_val_losses)\n        val_preds = torch.cat(epoch_val_preds)\n        val_targets = torch.cat(epoch_val_targets)\n        val_r2 = custom_r2_score(val_targets, val_preds)\n        \n        # Print metrics after each epoch\n        print(f\"🚀 Epoch [{epoch+1}/{num_epochs}] - \"\n              f\"Train Loss: {train_loss:.4f}, Train R²: {train_r2:.4f}, \"\n              f\"Val Loss: {val_loss:.4f}, Val R²: {val_r2:.4f}\")\n        \n        # Early Stopping check\n        if val_loss < best_val_loss:\n            best_val_loss = val_loss\n            patience_counter = 0\n            best_model_state = model.state_dict()\n            torch.save(best_model_state, f'best_attention_model_{epoch+1}.pth')\n            print(f'🏆 New best model saved with validation loss: {best_val_loss:.4f}')\n        else:\n            patience_counter += 1\n        \n        if patience_counter >= patience:\n            print(\"\\n📉 Early Stopping Triggered!\")\n            break\n        \n        # Save epoch metrics for later visualization\n        train_losses.append(train_loss)\n        val_losses.append(val_loss)\n        train_r2_scores.append(train_r2)\n        val_r2_scores.append(val_r2)\n    \n    # Final model loading\n    model.load_state_dict(best_model_state)\n    \n    # Training summary\n    total_time = time.time() - start_time\n    print(\"\\n🏁 Training Complete!\")\n    \n    # Plot training results\n    plt.figure(figsize=(12, 5))\n    plt.subplot(1, 2, 1)\n    plt.plot(train_losses, label='Train Loss')\n    plt.plot(val_losses, label='Val Loss')\n    plt.legend()\n    plt.title(\"Loss vs Epochs\")\n    \n    plt.subplot(1, 2, 2)\n    plt.plot(train_r2_scores, label='Train R²')\n    plt.plot(val_r2_scores, label='Val R²')\n    plt.legend()\n    plt.title(\"R² vs Epochs\")\n    \n    plt.show()\n\n# Example usage for training the model\nif __name__ == \"__main__\":\n    # Generate random data for example usage\n    batch_size = 128\n    seq_len = 10\n    input_dim = 79  # Prime number input dimension\n    print('📂 Loading dataset...')\n    df = pl.read_parquet(\"//kaggle/input/jane-street-real-time-market-data-forecasting/train.parquet\").filter(pl.col(\"partition_id\")>=4).fill_null(3)\n    \n    # Separate features and target\n    X = df[[f'feature_{i:02d}' for i in range(79)]+[\"symbol_id\",\"time_id\"]].to_numpy()\n    y = df[\"responder_6\"].to_numpy()\n    print(\"Dataset loaded.\")\n    del df\n    device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n    \n    # Train the model\n    train_hybrid_model(X, y, device)\n\n```",
              "votes": 7
            },
            {
              "id": 3060932,
              "postDate": "2024-12-02T09:11:27.283Z",
              "content": "<p>Thanks.</p>\n<p>Using an almost identical setup as well</p>",
              "rawMarkdown": "Thanks.\n\nUsing an almost identical setup as well"
            },
            {
              "id": 3060938,
              "postDate": "2024-12-02T09:18:41.377Z",
              "content": "<p>Oh, 1 more question: Does kaggle automatically let the models use all GPUs in parallel or do I need to account for that in my code</p>",
              "rawMarkdown": "Oh, 1 more question: Does kaggle automatically let the models use all GPUs in parallel or do I need to account for that in my code"
            },
            {
              "id": 3060943,
              "postDate": "2024-12-02T09:29:39Z",
              "content": "<p>You need to write the code for that, </p>\n<p>for ref:<br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nn.DataParallel.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.DataParallel.html</a></p>",
              "rawMarkdown": "You need to write the code for that, \n\nfor ref:\nhttps://pytorch.org/docs/stable/generated/torch.nn.DataParallel.html",
              "votes": 1
            },
            {
              "id": 3061324,
              "postDate": "2024-12-02T14:27:11.137Z",
              "content": "<blockquote>\n  <p>yeah training time is a problem on single or even on double t4 gpu's, use L4 gpus from other competitions,</p>\n</blockquote>\n<p>No. Absolutely no and please no. The L4 are limited for a specific competition (AIMO) for a reason, don't abuse Kaggle trust. If you need bigger compute, use TPU. It's far more than enough for this competition. The L4 is needed only for LLM comp' since not all pretrained models run on TPU, for this conpetition there is absolutely no justification for abusing the L4 when we have TPU. <a href=\"https://www.kaggle.com/zeroday19\" target=\"_blank\">@zeroday19</a> </p>\n<blockquote>\n  <blockquote>\n    <p>But also I still had trouble running 6 partitions through an MLP, 12 hours and it still didn't finish training 100 epochs, only trained 11 epochs</p>\n  </blockquote>\n</blockquote>\n<p>I can't tell for sure, but I suspect anything that is not online-training based would be useless. You better understand how to perform your training under the constraints of p100/t4 (NOT L4!) since this is what will be available to you for online training. Good luck. <a href=\"https://www.kaggle.com/pearsejim01\" target=\"_blank\">@pearsejim01</a> </p>",
              "rawMarkdown": ">yeah training time is a problem on single or even on double t4 gpu's, use L4 gpus from other competitions,\n\nNo. Absolutely no and please no. The L4 are limited for a specific competition (AIMO) for a reason, don't abuse Kaggle trust. If you need bigger compute, use TPU. It's far more than enough for this competition. The L4 is needed only for LLM comp' since not all pretrained models run on TPU, for this conpetition there is absolutely no justification for abusing the L4 when we have TPU. @zeroday19 \n\n>>But also I still had trouble running 6 partitions through an MLP, 12 hours and it still didn't finish training 100 epochs, only trained 11 epochs\n\nI can't tell for sure, but I suspect anything that is not online-training based would be useless. You better understand how to perform your training under the constraints of p100/t4 (NOT L4!) since this is what will be available to you for online training. Good luck. @pearsejim01 "
            },
            {
              "id": 3061389,
              "postDate": "2024-12-02T15:13:52.403Z",
              "content": "<p>I'm not very familiar with TPUs but I just tried to use them and the kaggle environment is somewhat different. It now doesn't have polars, and some other modules that come with kaggle env installed and I cant seem to install them.</p>",
              "rawMarkdown": "I'm not very familiar with TPUs but I just tried to use them and the kaggle environment is somewhat different. It now doesn't have polars, and some other modules that come with kaggle env installed and I cant seem to install them."
            },
            {
              "id": 3061395,
              "postDate": "2024-12-02T15:30:50.940Z",
              "content": "<p>Make sure that the notebook have internet connection enabled and then !pip install whatever you need. Tell me what you can't install and I will take a look.</p>",
              "rawMarkdown": "Make sure that the notebook have internet connection enabled and then !pip install whatever you need. Tell me what you can't install and I will take a look."
            },
            {
              "id": 3061404,
              "postDate": "2024-12-02T15:44:32.297Z",
              "content": "<p>I did enable internet connection. It still didnt install polars, lightgbm,catboost and xgboost</p>",
              "rawMarkdown": "I did enable internet connection. It still didnt install polars, lightgbm,catboost and xgboost"
            },
            {
              "id": 3061417,
              "postDate": "2024-12-02T15:52:22.957Z",
              "content": "<p>What batch size are you using? Try increasing it if you're on CPU</p>",
              "rawMarkdown": "What batch size are you using? Try increasing it if you're on CPU"
            },
            {
              "id": 3061448,
              "postDate": "2024-12-02T16:30:18.587Z",
              "content": "<p><a href=\"https://www.kaggle.com/pearsejim01\" target=\"_blank\">@pearsejim01</a> <a href=\"https://www.kaggle.com/code/shlomoron/tpu-install-polars\" target=\"_blank\">Here is polars on TPU</a>, works like a charm.  <br>\nlightgbm/catboost/xgboost won't work on TPU, it's for neural networks. So use GPU. But please no L4 unless it's for AIMO.</p>",
              "rawMarkdown": "@pearsejim01 [Here is polars on TPU](https://www.kaggle.com/code/shlomoron/tpu-install-polars), works like a charm.  \nlightgbm/catboost/xgboost won't work on TPU, it's for neural networks. So use GPU. But please no L4 unless it's for AIMO."
            },
            {
              "id": 3061452,
              "postDate": "2024-12-02T16:34:23.370Z",
              "content": "<p>Yeah I'm not using the GBDTs on TPU, they were just stuff I imported before. But polars still isn't working on my side.</p>\n<p>Also why won't it import the GBDTs on the notebook that enables TPU? Couldn't it just have the GBDTs use CPU while NNs use the TPU. </p>",
              "rawMarkdown": "Yeah I'm not using the GBDTs on TPU, they were just stuff I imported before. But polars still isn't working on my side.\n\nAlso why won't it import the GBDTs on the notebook that enables TPU? Couldn't it just have the GBDTs use CPU while NNs use the TPU. "
            },
            {
              "id": 3061456,
              "postDate": "2024-12-02T16:37:59.553Z",
              "content": "<p>Wait, it's working now, not sure what happened before, I ran the exact same notebook without change (it had pip install polars and then import polars), yet it brought errors just an hour ago but working smoothly now.</p>",
              "rawMarkdown": "Wait, it's working now, not sure what happened before, I ran the exact same notebook without change (it had pip install polars and then import polars), yet it brought errors just an hour ago but working smoothly now."
            },
            {
              "id": 3061472,
              "postDate": "2024-12-02T16:53:53.557Z",
              "content": "<p>Just remember that for submission you can only use p100/t4. TPU/L4 is really overkill considering that whatever you gonna do you will also want to be able to do it online</p>",
              "rawMarkdown": "Just remember that for submission you can only use p100/t4. TPU/L4 is really overkill considering that whatever you gonna do you will also want to be able to do it online"
            },
            {
              "id": 3062149,
              "postDate": "2024-12-03T10:35:39.050Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <a href=\"https://www.kaggle.com/pearsejim01\" target=\"_blank\">@pearsejim01</a> </p>\n<p>Apologies if I got the understanding from your above discussion wrong, but <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> you are advising we should write our training code such that it can train on both the training data and online test data using the provided p100/t4 right? Since for online training we only have access to that.</p>\n<p>However isnt the test data for online training much smaller than the given training data we have now? So wouldnt have 2 sets of code - one for the test data on kaggle gpus and one for the training data on an external gpu provider work too? Or am i missing something</p>",
              "rawMarkdown": "@shlomoron @pearsejim01 \n\nApologies if I got the understanding from your above discussion wrong, but @shlomoron you are advising we should write our training code such that it can train on both the training data and online test data using the provided p100/t4 right? Since for online training we only have access to that.\n\nHowever isnt the test data for online training much smaller than the given training data we have now? So wouldnt have 2 sets of code - one for the test data on kaggle gpus and one for the training data on an external gpu provider work too? Or am i missing something",
              "votes": 1
            },
            {
              "id": 3062156,
              "postDate": "2024-12-03T10:39:36.733Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3062158,
              "postDate": "2024-12-03T10:48:06.713Z",
              "content": "<p><a href=\"https://www.kaggle.com/junhaochan\" target=\"_blank\">@junhaochan</a> You are right! It's my mistake. I thought the full test set would be the size of the training set. Now I see it would only be the size of the current test set, which is much smaller. Well, this is why I like to discuss on the forum. Otherwise, I might not have noticed my mistake until the end! lol.<br>\nBut still, please don't use L4…</p>",
              "rawMarkdown": "@junhaochan You are right! It's my mistake. I thought the full test set would be the size of the training set. Now I see it would only be the size of the current test set, which is much smaller. Well, this is why I like to discuss on the forum. Otherwise, I might not have noticed my mistake until the end! lol.\nBut still, please don't use L4...",
              "votes": 1
            },
            {
              "id": 3062162,
              "postDate": "2024-12-03T10:53:15.640Z",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> </p>\n<p>ah wont be using L4, thanks for the speedy reply and clarification, much appreciated!</p>",
              "rawMarkdown": "@shlomoron \n\nah wont be using L4, thanks for the speedy reply and clarification, much appreciated!"
            }
          ]
        }
      ]
    },
    {
      "id": 3060530,
      "postDate": "2024-12-01T20:29:31.633Z",
      "content": "<p>Is there anyone here that has figured out a way to train all or most of the data on kaggle for Neural Networks? If so would you be willing to share your method?</p>",
      "rawMarkdown": "Is there anyone here that has figured out a way to train all or most of the data on kaggle for Neural Networks? If so would you be willing to share your method?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3060858,
      "author_name": "Abhi",
      "author_url": "",
      "post_date": "2024-12-02T07:09:34.933000",
      "content": "<p>its possible, i train all my model on data from partition 4 to 9, and it hardly takes 16 Gib of ram. use GPU, and load data with polars.</p>\n<p>In my tests pandas is not able to handle the data efficiently as polars is able to, and dont do splits with sklearn functions, do it manually, with numpy.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3060868,
          "author_name": "Seqaeon",
          "author_url": "",
          "post_date": "2024-12-02T07:21:28.283000",
          "content": "<p>I was talking more like 8-10 out of the partition. But also I still had trouble running 6 partitions through an MLP, 12 hours and it still didn't finish training 100 epochs, only trained  11 epochs. I load the data with polars but use pandas for some processing operations.</p>\n<p>Anyways, could you please share your template for it. You can remove the model or replace it with a simple MLP one, just need to see how you handle the data as it goes through the model. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3060877,
              "author_name": "Abhi",
              "author_url": "",
              "post_date": "2024-12-02T07:44:51.873000",
              "content": "<p>yeah training time is a problem on single or even on double t4 gpu's</p>\n<p>to look cool i used the emojis !!</p>\n<pre><code> torch\n torch.nn  nn\n torch.optim  optim\n torch.nn.functional  F\n math\n time\n numpy  np\n matplotlib.pyplot  plt\n\n\n\n ():\n    ss_total = ((y_true - y_true.mean())**).()\n    ss_residual = ((y_true - y_pred)**).()\n      - ss_residual / ss_total\n torch.utils.data  DataLoader, TensorDataset, Subset\n\n ():\n    \n    ()\n    indices = np.arange(X.shape[])  \n    \n\n    \n    train_size = ( * (indices))  \n    val_size = (indices) - train_size  \n\n    train_indices = indices[:train_size]\n    val_indices = indices[train_size:]\n\n    \n    ()\n    X_tensor = torch.FloatTensor(X).to(device)\n    y_tensor = torch.FloatTensor(y).to(device)\n     X,y\n    dataset = TensorDataset(X_tensor, y_tensor)\n    ()\n    \n    train_dataset = Subset(dataset, train_indices)\n    val_dataset = Subset(dataset, val_indices)\n\n    \n    train_loader = DataLoader(train_dataset, batch_size=batch_size, shuffle=)\n    val_loader = DataLoader(val_dataset, batch_size=batch_size, shuffle=)\n\n    \n    ()\n    model= NMDL(,, ) \n\n    \n    criterion = nn.MSELoss()\n    optimizer = optim.Adam(model.parameters(), lr=)\n    \n    best_val_loss = ()\n    patience_counter = \n    best_model_state = \n\n    \n    ()\n    start_time = time.time()\n\n    train_losses, val_losses = [], []\n    train_r2_scores, val_r2_scores = [], []\n\n     epoch  (num_epochs):\n        \n        ()\n        model.train()\n        epoch_train_losses = []\n        epoch_train_preds = []\n        epoch_train_targets = []\n\n         batch_X, batch_y  train_loader:\n            optimizer.zero_grad()\n\n            \n            train_pred = model(batch_X).squeeze()\n            loss = criterion(train_pred, batch_y)\n\n            \n            loss.backward()\n            optimizer.step()\n\n            \n            epoch_train_losses.append(loss.item())\n            epoch_train_preds.append(train_pred.detach())\n            epoch_train_targets.append(batch_y)\n\n        \n        train_loss = np.mean(epoch_train_losses)\n        train_preds = torch.cat(epoch_train_preds)\n        train_targets = torch.cat(epoch_train_targets)\n        train_r2 = custom_r2_score(train_targets, train_preds)\n\n        \n        ()\n        model.()\n        epoch_val_losses = []\n        epoch_val_preds = []\n        epoch_val_targets = []\n\n         torch.no_grad():\n             batch_X, batch_y  val_loader:\n                val_pred = model(batch_X).squeeze()\n                val_loss = criterion(val_pred, batch_y)\n\n                \n                epoch_val_losses.append(val_loss.item())\n                epoch_val_preds.append(val_pred)\n                epoch_val_targets.append(batch_y)\n\n        \n        val_loss = np.mean(epoch_val_losses)\n        val_preds = torch.cat(epoch_val_preds)\n        val_targets = torch.cat(epoch_val_targets)\n        val_r2 = custom_r2_score(val_targets, val_preds)\n\n        \n        (\n              \n              )\n\n        \n         val_loss &lt; best_val_loss:\n            best_val_loss = val_loss\n            patience_counter = \n            best_model_state = model.state_dict()\n            torch.save(best_model_state, )\n            ()\n        :\n            patience_counter += \n\n         patience_counter &gt;= patience:\n            ()\n            \n\n        \n        train_losses.append(train_loss)\n        val_losses.append(val_loss)\n        train_r2_scores.append(train_r2)\n        val_r2_scores.append(val_r2)\n\n    \n    model.load_state_dict(best_model_state)\n\n    \n    total_time = time.time() - start_time\n    ()\n\n    \n    plt.figure(figsize=(, ))\n    plt.subplot(, , )\n    plt.plot(train_losses, label=)\n    plt.plot(val_losses, label=)\n    plt.legend()\n    plt.title()\n\n    plt.subplot(, , )\n    plt.plot(train_r2_scores, label=)\n    plt.plot(val_r2_scores, label=)\n    plt.legend()\n    plt.title()\n\n    plt.show()\n\n\n __name__ == :\n    \n    batch_size = \n    seq_len = \n    input_dim =   \n    ()\n    df = pl.read_parquet().(pl.col()&gt;=).fill_null()\n\n    \n    X = df[[  i  ()]+[,]].to_numpy()\n    y = df[].to_numpy()\n    ()\n     df\n    device = torch.device(  torch.cuda.is_available()  )\n\n    \n    train_hybrid_model(X, y, device)\n</code></pre>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 3060932,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-12-02T09:11:27.283000",
              "content": "<p>Thanks.</p>\n<p>Using an almost identical setup as well</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3060938,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-12-02T09:18:41.377000",
              "content": "<p>Oh, 1 more question: Does kaggle automatically let the models use all GPUs in parallel or do I need to account for that in my code</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3060943,
              "author_name": "Abhi",
              "author_url": "",
              "post_date": "2024-12-02T09:29:39",
              "content": "<p>You need to write the code for that, </p>\n<p>for ref:<br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nn.DataParallel.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.DataParallel.html</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3061324,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-12-02T14:27:11.137000",
              "content": "<blockquote>\n  <p>yeah training time is a problem on single or even on double t4 gpu's, use L4 gpus from other competitions,</p>\n</blockquote>\n<p>No. Absolutely no and please no. The L4 are limited for a specific competition (AIMO) for a reason, don't abuse Kaggle trust. If you need bigger compute, use TPU. It's far more than enough for this competition. The L4 is needed only for LLM comp' since not all pretrained models run on TPU, for this conpetition there is absolutely no justification for abusing the L4 when we have TPU. <a href=\"https://www.kaggle.com/zeroday19\" target=\"_blank\">@zeroday19</a> </p>\n<blockquote>\n  <blockquote>\n    <p>But also I still had trouble running 6 partitions through an MLP, 12 hours and it still didn't finish training 100 epochs, only trained 11 epochs</p>\n  </blockquote>\n</blockquote>\n<p>I can't tell for sure, but I suspect anything that is not online-training based would be useless. You better understand how to perform your training under the constraints of p100/t4 (NOT L4!) since this is what will be available to you for online training. Good luck. <a href=\"https://www.kaggle.com/pearsejim01\" target=\"_blank\">@pearsejim01</a> </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061389,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-12-02T15:13:52.403000",
              "content": "<p>I'm not very familiar with TPUs but I just tried to use them and the kaggle environment is somewhat different. It now doesn't have polars, and some other modules that come with kaggle env installed and I cant seem to install them.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061395,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-12-02T15:30:50.940000",
              "content": "<p>Make sure that the notebook have internet connection enabled and then !pip install whatever you need. Tell me what you can't install and I will take a look.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061404,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-12-02T15:44:32.297000",
              "content": "<p>I did enable internet connection. It still didnt install polars, lightgbm,catboost and xgboost</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061417,
              "author_name": "Lu Bin Liu",
              "author_url": "",
              "post_date": "2024-12-02T15:52:22.957000",
              "content": "<p>What batch size are you using? Try increasing it if you're on CPU</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061448,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-12-02T16:30:18.587000",
              "content": "<p><a href=\"https://www.kaggle.com/pearsejim01\" target=\"_blank\">@pearsejim01</a> <a href=\"https://www.kaggle.com/code/shlomoron/tpu-install-polars\" target=\"_blank\">Here is polars on TPU</a>, works like a charm.  <br>\nlightgbm/catboost/xgboost won't work on TPU, it's for neural networks. So use GPU. But please no L4 unless it's for AIMO.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061452,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-12-02T16:34:23.370000",
              "content": "<p>Yeah I'm not using the GBDTs on TPU, they were just stuff I imported before. But polars still isn't working on my side.</p>\n<p>Also why won't it import the GBDTs on the notebook that enables TPU? Couldn't it just have the GBDTs use CPU while NNs use the TPU. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061456,
              "author_name": "Seqaeon",
              "author_url": "",
              "post_date": "2024-12-02T16:37:59.553000",
              "content": "<p>Wait, it's working now, not sure what happened before, I ran the exact same notebook without change (it had pip install polars and then import polars), yet it brought errors just an hour ago but working smoothly now.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3061472,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-12-02T16:53:53.557000",
              "content": "<p>Just remember that for submission you can only use p100/t4. TPU/L4 is really overkill considering that whatever you gonna do you will also want to be able to do it online</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3062149,
              "author_name": "just another tuesday",
              "author_url": "",
              "post_date": "2024-12-03T10:35:39.050000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> <a href=\"https://www.kaggle.com/pearsejim01\" target=\"_blank\">@pearsejim01</a> </p>\n<p>Apologies if I got the understanding from your above discussion wrong, but <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> you are advising we should write our training code such that it can train on both the training data and online test data using the provided p100/t4 right? Since for online training we only have access to that.</p>\n<p>However isnt the test data for online training much smaller than the given training data we have now? So wouldnt have 2 sets of code - one for the test data on kaggle gpus and one for the training data on an external gpu provider work too? Or am i missing something</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3062156,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-12-03T10:39:36.733000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3062158,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-12-03T10:48:06.713000",
              "content": "<p><a href=\"https://www.kaggle.com/junhaochan\" target=\"_blank\">@junhaochan</a> You are right! It's my mistake. I thought the full test set would be the size of the training set. Now I see it would only be the size of the current test set, which is much smaller. Well, this is why I like to discuss on the forum. Otherwise, I might not have noticed my mistake until the end! lol.<br>\nBut still, please don't use L4…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3062162,
              "author_name": "just another tuesday",
              "author_url": "",
              "post_date": "2024-12-03T10:53:15.640000",
              "content": "<p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> </p>\n<p>ah wont be using L4, thanks for the speedy reply and clarification, much appreciated!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3060858": "its possible, i train all my model on data from partition 4 to 9, and it hardly takes 16 Gib of ram. use GPU, and load data with polars.\n\nIn my tests pandas is not able to handle the data efficiently as polars is able to, and dont do splits with sklearn functions, do it manually, with numpy.\n \n",
    "3060530": "Is there anyone here that has figured out a way to train all or most of the data on kaggle for Neural Networks? If so would you be willing to share your method?"
  }
}