{
  "id": 535509,
  "title": "Can someone help me with this efficientnet training error?",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/535509",
  "author_name": "",
  "post_date": "2024-09-22T13:33:05.318095100Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.</p>\n<p>I tried to use <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline\" target=\"_blank\">this public notebook</a> to repeat the training process, but I found that in the training cycle, there would be a cpu memory explosion (not a gpu memory explosion), I tried to reduce n_workers to 8, this change just delayed the time point of cpu memory explosion, I wonder whether there is a memory leak in this code or some unobstructed place. But I didn't find the possible problem, so I would like to ask if anyone has encountered this problem, or can help me solve it? Thank you very much for your help!</p>\n<p>It could be the following code:</p>\n<pre><code>autocast = torch.cuda.amp.autocast(enabled=USE_AMP, dtype=torch.half) \nscaler = torch.cuda.amp.GradScaler(enabled=USE_AMP, init_scale=)\n\nskf = KFold(n_splits=N_FOLDS, shuffle=, random_state=SEED)\n fold, (trn_idx, val_idx)  (skf.split(((df)))):\n    (*)\n    ()\n    (*)\n    ((trn_idx), (val_idx))\n    df_train = df.iloc[trn_idx]\n    df_valid = df.iloc[val_idx]\n\n    train_ds = RSNA24Dataset(df_train, phase=, transform=transforms_train)\n    train_dl = DataLoader(\n                train_ds,\n                batch_size=BATCH_SIZE,\n                shuffle=,\n                pin_memory=,\n                drop_last=,\n                num_workers=N_WORKERS\n                )\n\n    valid_ds = RSNA24Dataset(df_valid, phase=, transform=transforms_val)\n    valid_dl = DataLoader(\n                valid_ds,\n                batch_size=BATCH_SIZE*,\n                shuffle=,\n                pin_memory=,\n                drop_last=,\n                num_workers=N_WORKERS\n                )\n\n    model = RSNA24Model(MODEL_NAME, IN_CHANS, N_CLASSES, pretrained=)\n    model.to(device)\n\n    optimizer = AdamW(model.parameters(), lr=LR, weight_decay=WD)\n\n    warmup_steps = EPOCHS/ * (train_dl) // GRAD_ACC\n    num_total_steps = EPOCHS * (train_dl) // GRAD_ACC\n    num_cycles = \n    scheduler = get_cosine_schedule_with_warmup(optimizer,\n                                                num_warmup_steps=warmup_steps,\n                                                num_training_steps=num_total_steps,\n                                                num_cycles=num_cycles)\n\n    weights = torch.tensor([, , ])\n    criterion = nn.CrossEntropyLoss(weight=weights.to(device))\n    criterion2 = nn.CrossEntropyLoss(weight=weights)\n\n    best_loss = \n    best_wll = \n    es_step = \n\n     epoch  (, EPOCHS+):\n        ()\n        model.train()\n        total_loss = \n         tqdm(train_dl, leave=)  pbar:\n            optimizer.zero_grad()\n             idx, (x, t)  (pbar):  \n                x = x.to(device)\n                t = t.to(device)\n\n                 autocast:\n                    loss = \n                    y = model(x)\n                     col  (N_LABELS):\n                        pred = y[:,col*:col*+]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / N_LABELS\n\n                    total_loss += loss.item()\n                     GRAD_ACC &gt; :\n                        loss = loss / GRAD_ACC\n\n                  math.isfinite(loss):\n                    ()\n                    sys.exit()\n\n                pbar.set_postfix(\n                    OrderedDict(\n                        loss=,\n                        lr=\n                    )\n                )\n                scaler.scale(loss).backward()\n\n                torch.nn.utils.clip_grad_norm_(model.parameters(), MAX_GRAD_NORM  )\n\n                 (idx + ) % GRAD_ACC == :\n                    scaler.step(optimizer)\n                    scaler.update()\n                    optimizer.zero_grad()\n                     scheduler   :\n                        scheduler.step()                    \n\n        train_loss = total_loss/(train_dl)\n        ()\n\n        total_loss = \n        y_preds = []\n        labels = []\n\n        model.()\n         tqdm(valid_dl, leave=)  pbar:\n             torch.no_grad():\n                 idx, (x, t)  (pbar):\n\n                    x = x.to(device)\n                    t = t.to(device)\n\n                     autocast:\n                        loss = \n                        loss_ema = \n                        y = model(x)\n                         col  (N_LABELS):\n                            pred = y[:,col*:col*+]\n                            gt = t[:,col]\n\n                            loss = loss + criterion(pred, gt) / N_LABELS\n                            y_pred = pred.()\n                            y_preds.append(y_pred.cpu())\n                            labels.append(gt.cpu())\n\n                        total_loss += loss.item()   \n\n        val_loss = total_loss/(valid_dl)\n</code></pre>",
  "messages": [
    {
      "id": "2995622",
      "postDate": "09/22/2024 13:33:05",
      "content": "<p>Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.</p>\n<p>I tried to use <a href=\"https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline\" target=\"_blank\">this public notebook</a> to repeat the training process, but I found that in the training cycle, there would be a cpu memory explosion (not a gpu memory explosion), I tried to reduce n_workers to 8, this change just delayed the time point of cpu memory explosion, I wonder whether there is a memory leak in this code or some unobstructed place. But I didn't find the possible problem, so I would like to ask if anyone has encountered this problem, or can help me solve it? Thank you very much for your help!</p>\n<p>It could be the following code:</p>\n<pre><code>autocast = torch.cuda.amp.autocast(enabled=USE_AMP, dtype=torch.half) \nscaler = torch.cuda.amp.GradScaler(enabled=USE_AMP, init_scale=)\n\nskf = KFold(n_splits=N_FOLDS, shuffle=, random_state=SEED)\n fold, (trn_idx, val_idx)  (skf.split(((df)))):\n    (*)\n    ()\n    (*)\n    ((trn_idx), (val_idx))\n    df_train = df.iloc[trn_idx]\n    df_valid = df.iloc[val_idx]\n\n    train_ds = RSNA24Dataset(df_train, phase=, transform=transforms_train)\n    train_dl = DataLoader(\n                train_ds,\n                batch_size=BATCH_SIZE,\n                shuffle=,\n                pin_memory=,\n                drop_last=,\n                num_workers=N_WORKERS\n                )\n\n    valid_ds = RSNA24Dataset(df_valid, phase=, transform=transforms_val)\n    valid_dl = DataLoader(\n                valid_ds,\n                batch_size=BATCH_SIZE*,\n                shuffle=,\n                pin_memory=,\n                drop_last=,\n                num_workers=N_WORKERS\n                )\n\n    model = RSNA24Model(MODEL_NAME, IN_CHANS, N_CLASSES, pretrained=)\n    model.to(device)\n\n    optimizer = AdamW(model.parameters(), lr=LR, weight_decay=WD)\n\n    warmup_steps = EPOCHS/ * (train_dl) // GRAD_ACC\n    num_total_steps = EPOCHS * (train_dl) // GRAD_ACC\n    num_cycles = \n    scheduler = get_cosine_schedule_with_warmup(optimizer,\n                                                num_warmup_steps=warmup_steps,\n                                                num_training_steps=num_total_steps,\n                                                num_cycles=num_cycles)\n\n    weights = torch.tensor([, , ])\n    criterion = nn.CrossEntropyLoss(weight=weights.to(device))\n    criterion2 = nn.CrossEntropyLoss(weight=weights)\n\n    best_loss = \n    best_wll = \n    es_step = \n\n     epoch  (, EPOCHS+):\n        ()\n        model.train()\n        total_loss = \n         tqdm(train_dl, leave=)  pbar:\n            optimizer.zero_grad()\n             idx, (x, t)  (pbar):  \n                x = x.to(device)\n                t = t.to(device)\n\n                 autocast:\n                    loss = \n                    y = model(x)\n                     col  (N_LABELS):\n                        pred = y[:,col*:col*+]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / N_LABELS\n\n                    total_loss += loss.item()\n                     GRAD_ACC &gt; :\n                        loss = loss / GRAD_ACC\n\n                  math.isfinite(loss):\n                    ()\n                    sys.exit()\n\n                pbar.set_postfix(\n                    OrderedDict(\n                        loss=,\n                        lr=\n                    )\n                )\n                scaler.scale(loss).backward()\n\n                torch.nn.utils.clip_grad_norm_(model.parameters(), MAX_GRAD_NORM  )\n\n                 (idx + ) % GRAD_ACC == :\n                    scaler.step(optimizer)\n                    scaler.update()\n                    optimizer.zero_grad()\n                     scheduler   :\n                        scheduler.step()                    \n\n        train_loss = total_loss/(train_dl)\n        ()\n\n        total_loss = \n        y_preds = []\n        labels = []\n\n        model.()\n         tqdm(valid_dl, leave=)  pbar:\n             torch.no_grad():\n                 idx, (x, t)  (pbar):\n\n                    x = x.to(device)\n                    t = t.to(device)\n\n                     autocast:\n                        loss = \n                        loss_ema = \n                        y = model(x)\n                         col  (N_LABELS):\n                            pred = y[:,col*:col*+]\n                            gt = t[:,col]\n\n                            loss = loss + criterion(pred, gt) / N_LABELS\n                            y_pred = pred.()\n                            y_preds.append(y_pred.cpu())\n                            labels.append(gt.cpu())\n\n                        total_loss += loss.item()   \n\n        val_loss = total_loss/(valid_dl)\n</code></pre>",
      "rawMarkdown": "Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.\n\nI tried to use [this public notebook](https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline) to repeat the training process, but I found that in the training cycle, there would be a cpu memory explosion (not a gpu memory explosion), I tried to reduce n_workers to 8, this change just delayed the time point of cpu memory explosion, I wonder whether there is a memory leak in this code or some unobstructed place. But I didn't find the possible problem, so I would like to ask if anyone has encountered this problem, or can help me solve it? Thank you very much for your help!\n\nIt could be the following code:\n```python\nautocast = torch.cuda.amp.autocast(enabled=USE_AMP, dtype=torch.half) # you can use with T4 gpu. or newer\nscaler = torch.cuda.amp.GradScaler(enabled=USE_AMP, init_scale=4096)\n\nskf = KFold(n_splits=N_FOLDS, shuffle=True, random_state=SEED)\nfor fold, (trn_idx, val_idx) in enumerate(skf.split(range(len(df)))):\n    print('#'*30)\n    print(f'start fold{fold}')\n    print('#'*30)\n    print(len(trn_idx), len(val_idx))\n    df_train = df.iloc[trn_idx]\n    df_valid = df.iloc[val_idx]\n\n    train_ds = RSNA24Dataset(df_train, phase='train', transform=transforms_train)\n    train_dl = DataLoader(\n                train_ds,\n                batch_size=BATCH_SIZE,\n                shuffle=True,\n                pin_memory=True,\n                drop_last=True,\n                num_workers=N_WORKERS\n                )\n\n    valid_ds = RSNA24Dataset(df_valid, phase='valid', transform=transforms_val)\n    valid_dl = DataLoader(\n                valid_ds,\n                batch_size=BATCH_SIZE*2,\n                shuffle=False,\n                pin_memory=True,\n                drop_last=False,\n                num_workers=N_WORKERS\n                )\n\n    model = RSNA24Model(MODEL_NAME, IN_CHANS, N_CLASSES, pretrained=False)\n    model.to(device)\n    \n    optimizer = AdamW(model.parameters(), lr=LR, weight_decay=WD)\n\n    warmup_steps = EPOCHS/10 * len(train_dl) // GRAD_ACC\n    num_total_steps = EPOCHS * len(train_dl) // GRAD_ACC\n    num_cycles = 0.475\n    scheduler = get_cosine_schedule_with_warmup(optimizer,\n                                                num_warmup_steps=warmup_steps,\n                                                num_training_steps=num_total_steps,\n                                                num_cycles=num_cycles)\n\n    weights = torch.tensor([1.0, 2.0, 4.0])\n    criterion = nn.CrossEntropyLoss(weight=weights.to(device))\n    criterion2 = nn.CrossEntropyLoss(weight=weights)\n\n    best_loss = 1.2\n    best_wll = 1.2\n    es_step = 0\n\n    for epoch in range(1, EPOCHS+1):\n        print(f'start epoch {epoch}')\n        model.train()\n        total_loss = 0\n        with tqdm(train_dl, leave=True) as pbar:\n            optimizer.zero_grad()\n            for idx, (x, t) in enumerate(pbar):  \n                x = x.to(device)\n                t = t.to(device)\n                \n                with autocast:\n                    loss = 0\n                    y = model(x)\n                    for col in range(N_LABELS):\n                        pred = y[:,col*3:col*3+3]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / N_LABELS\n                        \n                    total_loss += loss.item()\n                    if GRAD_ACC > 1:\n                        loss = loss / GRAD_ACC\n    \n                if not math.isfinite(loss):\n                    print(f\"Loss is {loss}, stopping training\")\n                    sys.exit(1)\n    \n                pbar.set_postfix(\n                    OrderedDict(\n                        loss=f'{loss.item()*GRAD_ACC:.6f}',\n                        lr=f'{optimizer.param_groups[0][\"lr\"]:.3e}'\n                    )\n                )\n                scaler.scale(loss).backward()\n\n                torch.nn.utils.clip_grad_norm_(model.parameters(), MAX_GRAD_NORM or 1e9)\n                \n                if (idx + 1) % GRAD_ACC == 0:\n                    scaler.step(optimizer)\n                    scaler.update()\n                    optimizer.zero_grad()\n                    if scheduler is not None:\n                        scheduler.step()                    \n    \n        train_loss = total_loss/len(train_dl)\n        print(f'train_loss:{train_loss:.6f}')\n\n        total_loss = 0\n        y_preds = []\n        labels = []\n        \n        model.eval()\n        with tqdm(valid_dl, leave=True) as pbar:\n            with torch.no_grad():\n                for idx, (x, t) in enumerate(pbar):\n                    \n                    x = x.to(device)\n                    t = t.to(device)\n                        \n                    with autocast:\n                        loss = 0\n                        loss_ema = 0\n                        y = model(x)\n                        for col in range(N_LABELS):\n                            pred = y[:,col*3:col*3+3]\n                            gt = t[:,col]\n \n                            loss = loss + criterion(pred, gt) / N_LABELS\n                            y_pred = pred.float()\n                            y_preds.append(y_pred.cpu())\n                            labels.append(gt.cpu())\n                        \n                        total_loss += loss.item()   \n    \n        val_loss = total_loss/len(valid_dl)\n\n```",
      "votes": null
    },
    {
      "id": "2995724",
      "postDate": "09/22/2024 14:47:57",
      "content": "<p>I believe setting <code>N_WORKERS</code> to <code>os.cpu_count()</code> addresses this problem. Let me know if that doesn't work, though; happy to try debug more.</p>",
      "rawMarkdown": "I believe setting `N_WORKERS` to `os.cpu_count()` addresses this problem. Let me know if that doesn't work, though; happy to try debug more.",
      "votes": null
    },
    {
      "id": "2995776",
      "postDate": "09/22/2024 15:16:56",
      "content": "<p>In fact, it's this <code>os.cpu_count()</code> that's causing the problem. The number of cores on my machine was greater than the number on kaggle's machine, and I only changed the number to 8. Later, I checked that the kaggle machine is 4 cores, so I changed it to 4 and it worked well. I'm still a little confused🥺, and may I ask if you have any relevant materials that I can learn? Thank you very much for your help</p>",
      "rawMarkdown": "In fact, it's this `os.cpu_count()` that's causing the problem. The number of cores on my machine was greater than the number on kaggle's machine, and I only changed the number to 8. Later, I checked that the kaggle machine is 4 cores, so I changed it to 4 and it worked well. I'm still a little confused🥺, and may I ask if you have any relevant materials that I can learn? Thank you very much for your help",
      "votes": null
    },
    {
      "id": "3004348",
      "postDate": "10/01/2024 18:02:42",
      "content": "<p>I would say you can not set the worker cups to the maximum cpus of your machine. You should always leave a few cpus for the other jobs of your local machine.</p>",
      "rawMarkdown": "I would say you can not set the worker cups to the maximum cpus of your machine. You should always leave a few cpus for the other jobs of your local machine.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2995724,
      "author_name": "bharatraghavan03",
      "author_url": "",
      "post_date": "09/22/2024 14:47:57",
      "content": "<p>I believe setting <code>N_WORKERS</code> to <code>os.cpu_count()</code> addresses this problem. Let me know if that doesn't work, though; happy to try debug more.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2995776,
          "author_name": "alannikos",
          "author_url": "",
          "post_date": "09/22/2024 15:16:56",
          "content": "<p>In fact, it's this <code>os.cpu_count()</code> that's causing the problem. The number of cores on my machine was greater than the number on kaggle's machine, and I only changed the number to 8. Later, I checked that the kaggle machine is 4 cores, so I changed it to 4 and it worked well. I'm still a little confused🥺, and may I ask if you have any relevant materials that I can learn? Thank you very much for your help</p>",
          "votes": null,
          "replies": [
            {
              "id": 3004348,
              "author_name": "denizcanelci",
              "author_url": "",
              "post_date": "10/01/2024 18:02:42",
              "content": "<p>I would say you can not set the worker cups to the maximum cpus of your machine. You should always leave a few cpus for the other jobs of your local machine.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2995622": "Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.\n\nI tried to use [this public notebook](https://www.kaggle.com/code/itsuki9180/rsna2024-lsdc-training-baseline) to repeat the training process, but I found that in the training cycle, there would be a cpu memory explosion (not a gpu memory explosion), I tried to reduce n_workers to 8, this change just delayed the time point of cpu memory explosion, I wonder whether there is a memory leak in this code or some unobstructed place. But I didn't find the possible problem, so I would like to ask if anyone has encountered this problem, or can help me solve it? Thank you very much for your help!\n\nIt could be the following code:\n```python\nautocast = torch.cuda.amp.autocast(enabled=USE_AMP, dtype=torch.half) # you can use with T4 gpu. or newer\nscaler = torch.cuda.amp.GradScaler(enabled=USE_AMP, init_scale=4096)\n\nskf = KFold(n_splits=N_FOLDS, shuffle=True, random_state=SEED)\nfor fold, (trn_idx, val_idx) in enumerate(skf.split(range(len(df)))):\n    print('#'*30)\n    print(f'start fold{fold}')\n    print('#'*30)\n    print(len(trn_idx), len(val_idx))\n    df_train = df.iloc[trn_idx]\n    df_valid = df.iloc[val_idx]\n\n    train_ds = RSNA24Dataset(df_train, phase='train', transform=transforms_train)\n    train_dl = DataLoader(\n                train_ds,\n                batch_size=BATCH_SIZE,\n                shuffle=True,\n                pin_memory=True,\n                drop_last=True,\n                num_workers=N_WORKERS\n                )\n\n    valid_ds = RSNA24Dataset(df_valid, phase='valid', transform=transforms_val)\n    valid_dl = DataLoader(\n                valid_ds,\n                batch_size=BATCH_SIZE*2,\n                shuffle=False,\n                pin_memory=True,\n                drop_last=False,\n                num_workers=N_WORKERS\n                )\n\n    model = RSNA24Model(MODEL_NAME, IN_CHANS, N_CLASSES, pretrained=False)\n    model.to(device)\n    \n    optimizer = AdamW(model.parameters(), lr=LR, weight_decay=WD)\n\n    warmup_steps = EPOCHS/10 * len(train_dl) // GRAD_ACC\n    num_total_steps = EPOCHS * len(train_dl) // GRAD_ACC\n    num_cycles = 0.475\n    scheduler = get_cosine_schedule_with_warmup(optimizer,\n                                                num_warmup_steps=warmup_steps,\n                                                num_training_steps=num_total_steps,\n                                                num_cycles=num_cycles)\n\n    weights = torch.tensor([1.0, 2.0, 4.0])\n    criterion = nn.CrossEntropyLoss(weight=weights.to(device))\n    criterion2 = nn.CrossEntropyLoss(weight=weights)\n\n    best_loss = 1.2\n    best_wll = 1.2\n    es_step = 0\n\n    for epoch in range(1, EPOCHS+1):\n        print(f'start epoch {epoch}')\n        model.train()\n        total_loss = 0\n        with tqdm(train_dl, leave=True) as pbar:\n            optimizer.zero_grad()\n            for idx, (x, t) in enumerate(pbar):  \n                x = x.to(device)\n                t = t.to(device)\n                \n                with autocast:\n                    loss = 0\n                    y = model(x)\n                    for col in range(N_LABELS):\n                        pred = y[:,col*3:col*3+3]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / N_LABELS\n                        \n                    total_loss += loss.item()\n                    if GRAD_ACC > 1:\n                        loss = loss / GRAD_ACC\n    \n                if not math.isfinite(loss):\n                    print(f\"Loss is {loss}, stopping training\")\n                    sys.exit(1)\n    \n                pbar.set_postfix(\n                    OrderedDict(\n                        loss=f'{loss.item()*GRAD_ACC:.6f}',\n                        lr=f'{optimizer.param_groups[0][\"lr\"]:.3e}'\n                    )\n                )\n                scaler.scale(loss).backward()\n\n                torch.nn.utils.clip_grad_norm_(model.parameters(), MAX_GRAD_NORM or 1e9)\n                \n                if (idx + 1) % GRAD_ACC == 0:\n                    scaler.step(optimizer)\n                    scaler.update()\n                    optimizer.zero_grad()\n                    if scheduler is not None:\n                        scheduler.step()                    \n    \n        train_loss = total_loss/len(train_dl)\n        print(f'train_loss:{train_loss:.6f}')\n\n        total_loss = 0\n        y_preds = []\n        labels = []\n        \n        model.eval()\n        with tqdm(valid_dl, leave=True) as pbar:\n            with torch.no_grad():\n                for idx, (x, t) in enumerate(pbar):\n                    \n                    x = x.to(device)\n                    t = t.to(device)\n                        \n                    with autocast:\n                        loss = 0\n                        loss_ema = 0\n                        y = model(x)\n                        for col in range(N_LABELS):\n                            pred = y[:,col*3:col*3+3]\n                            gt = t[:,col]\n \n                            loss = loss + criterion(pred, gt) / N_LABELS\n                            y_pred = pred.float()\n                            y_preds.append(y_pred.cpu())\n                            labels.append(gt.cpu())\n                        \n                        total_loss += loss.item()   \n    \n        val_loss = total_loss/len(valid_dl)\n\n```",
    "2995724": "I believe setting `N_WORKERS` to `os.cpu_count()` addresses this problem. Let me know if that doesn't work, though; happy to try debug more.",
    "2995776": "In fact, it's this `os.cpu_count()` that's causing the problem. The number of cores on my machine was greater than the number on kaggle's machine, and I only changed the number to 8. Later, I checked that the kaggle machine is 4 cores, so I changed it to 4 and it worked well. I'm still a little confused🥺, and may I ask if you have any relevant materials that I can learn? Thank you very much for your help",
    "3004348": "I would say you can not set the worker cups to the maximum cpus of your machine. You should always leave a few cpus for the other jobs of your local machine."
  },
  "source": "meta"
}