{
  "id": 515938,
  "title": "[Pytorch] Really Need Some HELP !!!",
  "url": "/competitions/leash-BELKA/discussion/515938",
  "author_name": "",
  "post_date": "2024-06-30T18:17:29.189664300Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>These day I just tried to implement the pytorch version of this amazing notebook:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook</a></p>\n<p>I extracted about 650M samples from the same dataset, but my train_loss always remained at about 0.16 and the val_loss remained at about 0.4, which is terrible. And of course, I got really bad score on public. <br>\nAfter 2 days' debugging, I still can't figure out what's going on. So I come here to seek some suggestions ~<br>\nHere are the key parts of my codes, can anyone give me some advice? Really appreciate it !!!</p>\n<p><code>class CFG: \n    preprocess = False\n    pretrain = False\n    batch_size = 2048\n    learning_rate = 1e-4\n    AdamW_decay = 1e-6\n    device = torch.device('cuda') if torch.cuda.is_available() else 'cpu'\n    folds = 10</code></p>\n<p>`class OneDCNN(torch.nn.Module):<br>\n    def <strong>init</strong>(self):<br>\n        super().<strong>init</strong>()</p>\n<pre><code>     = \n     = \n\n    .embedding = torch.nn.Embedding(\n        num_embeddings=, \n        embedding_dim=,\n        padding_idx=\n    )  \n\n    .model1 = torch.nn.Sequential(\n        torch.nn.Conv1d(, , ),  \n        torch.nn.ReLU(),\n        torch.nn.Conv1d(, , ),  \n        torch.nn.ReLU(),\n        torch.nn.Conv1d(, , ),  \n        torch.nn.AdaptiveMaxPool1d()  \n    )\n\n    .model2 = torch.nn.Sequential(\n        torch.nn.Linear(, ),\n        torch.nn.ReLU(),\n        torch.nn.Linear(, ),\n        torch.nn.ReLU(),\n        torch.nn.Linear(, )\n    )\n\n ():\n    x = x.long()\n    x = .embedding(x)  \n    x = x.permute(, , )  \n    x = .model1(x)\n    x = x.squeeze()  \n    x = .model2(x)\n     x</code></pre>\n<p>`class Trainer:<br>\n    def <strong>init</strong>(self, loaders, model):<br>\n        self.train_loader, self.val_loader = loaders<br>\n        self.model = model<br>\n        self.optim = torch.optim.AdamW(self.model.parameters(), lr=CFG.learning_rate, weight_decay=CFG.AdamW_decay)</p>\n<pre><code>    self.train_loss = []\n    self.val_loss = []\n\n ():\n    gc.collect()\n    torch.cuda.empty_cache()\n\n ():\n    running_loss = \n    progress = tqdm(self.train_loader, total=(self.train_loader))\n\n     i, (features, targets)  (progress):\n        self.optim.zero_grad()\n\n        targets_pred = self.model(features)\n        loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n        running_loss += loss.item()\n\n        loss.backward()\n        self.optim.step()\n\n    train_loss = running_loss / (self.train_loader)\n    self.train_loss.append(train_loss)\n\n\n ():\n    running_loss = \n    progress = tqdm(self.val_loader, total=(self.val_loader))\n\n     (features, targets)  progress:\n        targets_pred = self.model(features)\n        loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n        running_loss += loss.item()\n\n    val_loss = running_loss / (self.val_loader)\n    self.val_loss.append(val_loss)\n\n ():\n    progress = tqdm((, ), desc=)\n\n     epoch  progress:\n        progress.set_description()\n        self.train_one_epoch()\n        self.clear()\n\n        progress.set_description()\n        self.val_one_epoch()\n        self.clear()\n\n        ()\n        ()\n        ()\n\n ():\n    preds = []\n\n     features  test_loader:\n        targets_pred = self.model(features)\n        preds.append(targets_pred.detach().cpu())\n\n    preds = torch.concat(preds)\n     preds\n\n ():\n     self.model`\n</code></pre>",
  "messages": [
    {
      "id": "2897866",
      "postDate": "06/30/2024 18:17:29",
      "content": "<p>These day I just tried to implement the pytorch version of this amazing notebook:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook</a></p>\n<p>I extracted about 650M samples from the same dataset, but my train_loss always remained at about 0.16 and the val_loss remained at about 0.4, which is terrible. And of course, I got really bad score on public. <br>\nAfter 2 days' debugging, I still can't figure out what's going on. So I come here to seek some suggestions ~<br>\nHere are the key parts of my codes, can anyone give me some advice? Really appreciate it !!!</p>\n<p><code>class CFG: \n    preprocess = False\n    pretrain = False\n    batch_size = 2048\n    learning_rate = 1e-4\n    AdamW_decay = 1e-6\n    device = torch.device('cuda') if torch.cuda.is_available() else 'cpu'\n    folds = 10</code></p>\n<p>`class OneDCNN(torch.nn.Module):<br>\n    def <strong>init</strong>(self):<br>\n        super().<strong>init</strong>()</p>\n<pre><code>     = \n     = \n\n    .embedding = torch.nn.Embedding(\n        num_embeddings=, \n        embedding_dim=,\n        padding_idx=\n    )  \n\n    .model1 = torch.nn.Sequential(\n        torch.nn.Conv1d(, , ),  \n        torch.nn.ReLU(),\n        torch.nn.Conv1d(, , ),  \n        torch.nn.ReLU(),\n        torch.nn.Conv1d(, , ),  \n        torch.nn.AdaptiveMaxPool1d()  \n    )\n\n    .model2 = torch.nn.Sequential(\n        torch.nn.Linear(, ),\n        torch.nn.ReLU(),\n        torch.nn.Linear(, ),\n        torch.nn.ReLU(),\n        torch.nn.Linear(, )\n    )\n\n ():\n    x = x.long()\n    x = .embedding(x)  \n    x = x.permute(, , )  \n    x = .model1(x)\n    x = x.squeeze()  \n    x = .model2(x)\n     x</code></pre>\n<p>`class Trainer:<br>\n    def <strong>init</strong>(self, loaders, model):<br>\n        self.train_loader, self.val_loader = loaders<br>\n        self.model = model<br>\n        self.optim = torch.optim.AdamW(self.model.parameters(), lr=CFG.learning_rate, weight_decay=CFG.AdamW_decay)</p>\n<pre><code>    self.train_loss = []\n    self.val_loss = []\n\n ():\n    gc.collect()\n    torch.cuda.empty_cache()\n\n ():\n    running_loss = \n    progress = tqdm(self.train_loader, total=(self.train_loader))\n\n     i, (features, targets)  (progress):\n        self.optim.zero_grad()\n\n        targets_pred = self.model(features)\n        loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n        running_loss += loss.item()\n\n        loss.backward()\n        self.optim.step()\n\n    train_loss = running_loss / (self.train_loader)\n    self.train_loss.append(train_loss)\n\n\n ():\n    running_loss = \n    progress = tqdm(self.val_loader, total=(self.val_loader))\n\n     (features, targets)  progress:\n        targets_pred = self.model(features)\n        loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n        running_loss += loss.item()\n\n    val_loss = running_loss / (self.val_loader)\n    self.val_loss.append(val_loss)\n\n ():\n    progress = tqdm((, ), desc=)\n\n     epoch  progress:\n        progress.set_description()\n        self.train_one_epoch()\n        self.clear()\n\n        progress.set_description()\n        self.val_one_epoch()\n        self.clear()\n\n        ()\n        ()\n        ()\n\n ():\n    preds = []\n\n     features  test_loader:\n        targets_pred = self.model(features)\n        preds.append(targets_pred.detach().cpu())\n\n    preds = torch.concat(preds)\n     preds\n\n ():\n     self.model`\n</code></pre>",
      "rawMarkdown": "These day I just tried to implement the pytorch version of this amazing notebook:\n[https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook](url)\n\nI extracted about 650M samples from the same dataset, but my train_loss always remained at about 0.16 and the val_loss remained at about 0.4, which is terrible. And of course, I got really bad score on public. \nAfter 2 days' debugging, I still can't figure out what's going on. So I come here to seek some suggestions ~\nHere are the key parts of my codes, can anyone give me some advice? Really appreciate it !!!\n\n`class CFG: \n    preprocess = False\n    pretrain = False\n    batch_size = 2048\n    learning_rate = 1e-4\n    AdamW_decay = 1e-6\n    device = torch.device('cuda') if torch.cuda.is_available() else 'cpu'\n    folds = 10`\n\n`class OneDCNN(torch.nn.Module):\n    def __init__(self):\n        super().__init__()\n        \n        INPUT_DIM = 37\n        HIDDEN_DIM = 128\n        \n        self.embedding = torch.nn.Embedding(\n            num_embeddings=INPUT_DIM, \n            embedding_dim=HIDDEN_DIM,\n            padding_idx=0\n        )  # [BATCH_SIZE, 142] --> [BATCH_SIZE, 142, 128]\n        \n        self.model1 = torch.nn.Sequential(\n            torch.nn.Conv1d(128, 32, 3),  # [BATCH_SIZE, 128, 142] --> [BATCH_SIZE, 32, 140]\n            torch.nn.ReLU(),\n            torch.nn.Conv1d(32, 64, 3),  # [BATCH_SIZE, 32, 140] --> [BATCH_SIZE, 64, 138]\n            torch.nn.ReLU(),\n            torch.nn.Conv1d(64, 96, 3),  # [BATCH_SIZE, 64, 138] --> [BATCH_SIZE, 96, 136]\n            torch.nn.AdaptiveMaxPool1d(1)  # [BATCH_SIZE, 96, 136] --> [BATCH_SIZE, 96, 1]\n        )\n        \n        self.model2 = torch.nn.Sequential(\n            torch.nn.Linear(96, 1024),\n            torch.nn.ReLU(),\n            torch.nn.Linear(1024, 512),\n            torch.nn.ReLU(),\n            torch.nn.Linear(512, 3)\n        )\n    \n    def forward(self, x):\n        x = x.long()\n        x = self.embedding(x)  # [BATCH_SIZE, 142] --> [BATCH_SIZE, 142, 128]\n        x = x.permute(0, 2, 1)  # [BATCH_SIZE, 142, 128] --> [BATCH_SIZE, 128, 142]\n        x = self.model1(x)\n        x = x.squeeze(2)  # [BATCH_SIZE, 96, 1] --> [BATCH_SIZE, 96]\n        x = self.model2(x)\n        return x`\n\n`class Trainer:\n    def __init__(self, loaders, model):\n        self.train_loader, self.val_loader = loaders\n        self.model = model\n        self.optim = torch.optim.AdamW(self.model.parameters(), lr=CFG.learning_rate, weight_decay=CFG.AdamW_decay)\n        \n        self.train_loss = []\n        self.val_loss = []\n    \n    def clear(self):\n        gc.collect()\n        torch.cuda.empty_cache()\n    \n    def train_one_epoch(self):\n        running_loss = 0\n        progress = tqdm(self.train_loader, total=len(self.train_loader))\n        \n        for i, (features, targets) in enumerate(progress):\n            self.optim.zero_grad()\n            \n            targets_pred = self.model(features)\n            loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n            running_loss += loss.item()\n            \n            loss.backward()\n            self.optim.step()\n        \n        train_loss = running_loss / len(self.train_loader)\n        self.train_loss.append(train_loss)\n    \n    @torch.no_grad()\n    def val_one_epoch(self):\n        running_loss = 0\n        progress = tqdm(self.val_loader, total=len(self.val_loader))\n        \n        for (features, targets) in progress:\n            targets_pred = self.model(features)\n            loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n            running_loss += loss.item()\n            \n        val_loss = running_loss / len(self.val_loader)\n        self.val_loss.append(val_loss)\n    \n    def fit(self):\n        progress = tqdm(range(1, 2), desc='Training...')\n        \n        for epoch in progress:\n            progress.set_description(f\"EPOCH{epoch}/{1} | training...\")\n            self.train_one_epoch()\n            self.clear()\n        \n            progress.set_description(f\"EPOCH{epoch}/{1} | validating...\")\n            self.val_one_epoch()\n            self.clear()\n            \n            print(f\"{'-'*30} EPOCH {epoch} / {1} {'-'*30}\")\n            print(f\"train_loss: {self.train_loss[-1]}\")\n            print(f\"val_loss: {self.val_loss[-1]}\")\n            \n    def predict(self, test_loader):\n        preds = []\n        \n        for features in test_loader:\n            targets_pred = self.model(features)\n            preds.append(targets_pred.detach().cpu())\n        \n        preds = torch.concat(preds)\n        return preds\n    \n    def get_model(self):\n        return self.model`",
      "votes": null
    },
    {
      "id": "2898026",
      "postDate": "06/30/2024 21:47:34",
      "content": "<p>Okay, I see you're new to the PyTorch ecosystem. Let me give you some pointers to help you out.</p>\n<p>First of all, data preprocessing is crucial. Make sure to scale or normalize your features using something like scikit-learn's StandardScaler. This can help your model converge faster and more stably.</p>\n<p>Next, your learning rate of 1e-4 might be a bit too high for such a large dataset. Try lowering it to 1e-5 or even 1e-6 and see if that improves things.</p>\n<p>The large gap between your training and validation losses points to overfitting. Consider adding some regularization, like dropout layers or L2 regularization on the weights. Also, make sure your model's complexity is appropriate for your task - not too simple to underfit, but not too complex to overfit.</p>\n<p>Check if your dataset has class imbalance. If some classes have way more samples than others, your model might be biased. You can fix this by oversampling the minority classes, undersampling the majority classes, or adjusting class weights in the loss function.</p>\n<p>Make sure you're using the right loss function. For multi-class classification, use CrossEntropyLoss instead of binary cross-entropy.</p>\n<p>To make your life easier, use PyTorch Lightning. It's a lightweight PyTorch wrapper that handles a lot of the boring stuff for you. Here's how your model could look with Lightning:</p>\n<pre><code> ():\n     ():\n        ().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, , kernel_size=) \n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d()\n        self.fc = nn.Linear(, num_classes)\n\n     ():\n        x = self.embedding(x) \n        x = torch.einsum(, x)  \n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze()\n        x = self.fc(x)\n         x\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n         loss\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n\n     ():\n         torch.optim.AdamW(self.parameters(), lr=)\n</code></pre>\n<p>I snuck in an <a href=\"einsum\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.einsum.html</a> operation for the transpose in the forward method. Einsum is a super handy tool for tensor ops that can often be faster and cleaner than chaining together a bunch of separate PyTorch functions.</p>\n<p>Here's a minimum viable example:</p>\n<pre><code> torch\n torch  nn\n torch.utils.data  DataLoader, TensorDataset\n pytorch_lightning  LightningModule, Trainer\n\n\n ():\n     ():\n        ().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, , kernel_size=)\n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d()\n        self.fc = nn.Linear(, num_classes)\n\n     ():\n        x = self.embedding(x)\n        x = torch.einsum(, x)  \n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze()\n        x = self.fc(x)\n         x\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n         loss\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n\n     ():\n         torch.optim.AdamW(self.parameters(), lr=)\n\n\n\nX = torch.randint(, , (, ))\ny = torch.randint(, , (,))\n\n\ndataset = TensorDataset(X, y)\ntrain_loader = DataLoader(dataset, batch_size=)\nval_loader = DataLoader(dataset, batch_size=)\n\n\nmodel = SimpleOneDCNN()\ntrainer = Trainer(max_epochs=)\ntrainer.fit(model, train_loader, val_loader)\n</code></pre>\n<p><strong>One thing to note… don't focus too much on the actual <em>value</em> of the loss, just look for the loss to decrease in a steady way - watch for numerical stability above all else</strong></p>\n<p>Good luck :) </p>",
      "rawMarkdown": "Okay, I see you're new to the PyTorch ecosystem. Let me give you some pointers to help you out.\n\nFirst of all, data preprocessing is crucial. Make sure to scale or normalize your features using something like scikit-learn's StandardScaler. This can help your model converge faster and more stably.\n\nNext, your learning rate of 1e-4 might be a bit too high for such a large dataset. Try lowering it to 1e-5 or even 1e-6 and see if that improves things.\n\nThe large gap between your training and validation losses points to overfitting. Consider adding some regularization, like dropout layers or L2 regularization on the weights. Also, make sure your model's complexity is appropriate for your task - not too simple to underfit, but not too complex to overfit.\n\nCheck if your dataset has class imbalance. If some classes have way more samples than others, your model might be biased. You can fix this by oversampling the minority classes, undersampling the majority classes, or adjusting class weights in the loss function.\n\nMake sure you're using the right loss function. For multi-class classification, use CrossEntropyLoss instead of binary cross-entropy.\n\nTo make your life easier, use PyTorch Lightning. It's a lightweight PyTorch wrapper that handles a lot of the boring stuff for you. Here's how your model could look with Lightning:\n\n```python\nclass SimpleOneDCNN(LightningModule):\n    def __init__(self, input_dim=37, embedding_dim=128, num_classes=3):\n        super().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, 64, kernel_size=3) \n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d(1)\n        self.fc = nn.Linear(64, num_classes)\n        \n    def forward(self, x):\n        x = self.embedding(x) \n        x = torch.einsum(\"bse->bes\", x)  # transpose for Conv1d\n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze(2)\n        x = self.fc(x)\n        return x\n\n    def training_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('train_loss', loss)\n        return loss\n\n    def validation_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('val_loss', loss)\n\n    def configure_optimizers(self):\n        return torch.optim.AdamW(self.parameters(), lr=1e-5)\n```\n\nI snuck in an [https://pytorch.org/docs/stable/generated/torch.einsum.html](einsum) operation for the transpose in the forward method. Einsum is a super handy tool for tensor ops that can often be faster and cleaner than chaining together a bunch of separate PyTorch functions.\n\nHere's a minimum viable example:\n\n```python\nimport torch\nfrom torch import nn\nfrom torch.utils.data import DataLoader, TensorDataset\nfrom pytorch_lightning import LightningModule, Trainer\n\n\nclass SimpleOneDCNN(LightningModule):\n    def __init__(self, input_dim=37, embedding_dim=128, num_classes=3):\n        super().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, 64, kernel_size=3)\n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d(1)\n        self.fc = nn.Linear(64, num_classes)\n\n    def forward(self, x):\n        x = self.embedding(x)\n        x = torch.einsum(\"bse->bes\", x)  # transpose for Conv1d\n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze(2)\n        x = self.fc(x)\n        return x\n\n    def training_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('train_loss', loss)\n        return loss\n\n    def validation_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('val_loss', loss)\n\n    def configure_optimizers(self):\n        return torch.optim.AdamW(self.parameters(), lr=1e-3)\n\n\n# Example synthetic data\nX = torch.randint(0, 37, (1000, 142))\ny = torch.randint(0, 3, (1000,))\n\n# Create DataLoaders\ndataset = TensorDataset(X, y)\ntrain_loader = DataLoader(dataset, batch_size=32)\nval_loader = DataLoader(dataset, batch_size=32)\n\n# Train the model\nmodel = SimpleOneDCNN()\ntrainer = Trainer(max_epochs=5)\ntrainer.fit(model, train_loader, val_loader)\n```\n**One thing to note... don't focus too much on the actual *value* of the loss, just look for the loss to decrease in a steady way - watch for numerical stability above all else**\n\nGood luck :)",
      "votes": null
    },
    {
      "id": "2898499",
      "postDate": "07/01/2024 07:18:16",
      "content": "<p>Thanks a lot ! I will do a quick study on Lightning immediately</p>",
      "rawMarkdown": "Thanks a lot ! I will do a quick study on Lightning immediately",
      "votes": null
    },
    {
      "id": "2899918",
      "postDate": "07/01/2024 23:29:11",
      "content": "<p>Btw, feel free to take a look at my reimplementation of the 1dcnn notebook in PyTorch :)</p>\n<p><a href=\"https://www.kaggle.com/code/yyyu54/pytorch-version-belka-1dcnn-starter-with-all-data\" target=\"_blank\">Link</a></p>",
      "rawMarkdown": "Btw, feel free to take a look at my reimplementation of the 1dcnn notebook in PyTorch :)\n\n[Link](https://www.kaggle.com/code/yyyu54/pytorch-version-belka-1dcnn-starter-with-all-data)",
      "votes": null
    },
    {
      "id": "2900311",
      "postDate": "07/02/2024 07:34:57",
      "content": "<p>YESYES !! I have tried the Pytorch Lightning Module, and just follow your wonderful notebook to design my new 1dcnn model. Now I can reach 0.386 on the public !! Thanks for your work :) Upvote !!</p>",
      "rawMarkdown": "YESYES !! I have tried the Pytorch Lightning Module, and just follow your wonderful notebook to design my new 1dcnn model. Now I can reach 0.386 on the public !! Thanks for your work :) Upvote !!",
      "votes": null
    },
    {
      "id": "2900318",
      "postDate": "07/02/2024 07:38:46",
      "content": "<p>Actually, I have read your notebook a few days ago. But at that time, I was still unfamiliar with the Lightning Module. And now, I have found that it's much more convenient and powerful than the original one :) (By the way, you are 1st in this competition ! Omg. Amazing.</p>",
      "rawMarkdown": "Actually, I have read your notebook a few days ago. But at that time, I was still unfamiliar with the Lightning Module. And now, I have found that it's much more convenient and powerful than the original one :) (By the way, you are 1st in this competition ! Omg. Amazing.",
      "votes": null
    },
    {
      "id": "2900551",
      "postDate": "07/02/2024 10:56:55",
      "content": "<p>Glad it worked for you! Best of luck for the last couple of days! </p>",
      "rawMarkdown": "Glad it worked for you! Best of luck for the last couple of days!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2898026,
      "author_name": "mbanuna",
      "author_url": "",
      "post_date": "06/30/2024 21:47:34",
      "content": "<p>Okay, I see you're new to the PyTorch ecosystem. Let me give you some pointers to help you out.</p>\n<p>First of all, data preprocessing is crucial. Make sure to scale or normalize your features using something like scikit-learn's StandardScaler. This can help your model converge faster and more stably.</p>\n<p>Next, your learning rate of 1e-4 might be a bit too high for such a large dataset. Try lowering it to 1e-5 or even 1e-6 and see if that improves things.</p>\n<p>The large gap between your training and validation losses points to overfitting. Consider adding some regularization, like dropout layers or L2 regularization on the weights. Also, make sure your model's complexity is appropriate for your task - not too simple to underfit, but not too complex to overfit.</p>\n<p>Check if your dataset has class imbalance. If some classes have way more samples than others, your model might be biased. You can fix this by oversampling the minority classes, undersampling the majority classes, or adjusting class weights in the loss function.</p>\n<p>Make sure you're using the right loss function. For multi-class classification, use CrossEntropyLoss instead of binary cross-entropy.</p>\n<p>To make your life easier, use PyTorch Lightning. It's a lightweight PyTorch wrapper that handles a lot of the boring stuff for you. Here's how your model could look with Lightning:</p>\n<pre><code> ():\n     ():\n        ().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, , kernel_size=) \n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d()\n        self.fc = nn.Linear(, num_classes)\n\n     ():\n        x = self.embedding(x) \n        x = torch.einsum(, x)  \n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze()\n        x = self.fc(x)\n         x\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n         loss\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n\n     ():\n         torch.optim.AdamW(self.parameters(), lr=)\n</code></pre>\n<p>I snuck in an <a href=\"einsum\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.einsum.html</a> operation for the transpose in the forward method. Einsum is a super handy tool for tensor ops that can often be faster and cleaner than chaining together a bunch of separate PyTorch functions.</p>\n<p>Here's a minimum viable example:</p>\n<pre><code> torch\n torch  nn\n torch.utils.data  DataLoader, TensorDataset\n pytorch_lightning  LightningModule, Trainer\n\n\n ():\n     ():\n        ().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, , kernel_size=)\n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d()\n        self.fc = nn.Linear(, num_classes)\n\n     ():\n        x = self.embedding(x)\n        x = torch.einsum(, x)  \n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze()\n        x = self.fc(x)\n         x\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n         loss\n\n     ():\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log(, loss)\n\n     ():\n         torch.optim.AdamW(self.parameters(), lr=)\n\n\n\nX = torch.randint(, , (, ))\ny = torch.randint(, , (,))\n\n\ndataset = TensorDataset(X, y)\ntrain_loader = DataLoader(dataset, batch_size=)\nval_loader = DataLoader(dataset, batch_size=)\n\n\nmodel = SimpleOneDCNN()\ntrainer = Trainer(max_epochs=)\ntrainer.fit(model, train_loader, val_loader)\n</code></pre>\n<p><strong>One thing to note… don't focus too much on the actual <em>value</em> of the loss, just look for the loss to decrease in a steady way - watch for numerical stability above all else</strong></p>\n<p>Good luck :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 2898499,
          "author_name": "kalvinlt",
          "author_url": "",
          "post_date": "07/01/2024 07:18:16",
          "content": "<p>Thanks a lot ! I will do a quick study on Lightning immediately</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2899918,
      "author_name": "yyyu54",
      "author_url": "",
      "post_date": "07/01/2024 23:29:11",
      "content": "<p>Btw, feel free to take a look at my reimplementation of the 1dcnn notebook in PyTorch :)</p>\n<p><a href=\"https://www.kaggle.com/code/yyyu54/pytorch-version-belka-1dcnn-starter-with-all-data\" target=\"_blank\">Link</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2900311,
          "author_name": "kalvinlt",
          "author_url": "",
          "post_date": "07/02/2024 07:34:57",
          "content": "<p>YESYES !! I have tried the Pytorch Lightning Module, and just follow your wonderful notebook to design my new 1dcnn model. Now I can reach 0.386 on the public !! Thanks for your work :) Upvote !!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2900318,
          "author_name": "kalvinlt",
          "author_url": "",
          "post_date": "07/02/2024 07:38:46",
          "content": "<p>Actually, I have read your notebook a few days ago. But at that time, I was still unfamiliar with the Lightning Module. And now, I have found that it's much more convenient and powerful than the original one :) (By the way, you are 1st in this competition ! Omg. Amazing.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2900551,
              "author_name": "yyyu54",
              "author_url": "",
              "post_date": "07/02/2024 10:56:55",
              "content": "<p>Glad it worked for you! Best of luck for the last couple of days! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2897866": "These day I just tried to implement the pytorch version of this amazing notebook:\n[https://www.kaggle.com/code/ahmedelfazouan/belka-1dcnn-starter-with-all-data/notebook](url)\n\nI extracted about 650M samples from the same dataset, but my train_loss always remained at about 0.16 and the val_loss remained at about 0.4, which is terrible. And of course, I got really bad score on public. \nAfter 2 days' debugging, I still can't figure out what's going on. So I come here to seek some suggestions ~\nHere are the key parts of my codes, can anyone give me some advice? Really appreciate it !!!\n\n`class CFG: \n    preprocess = False\n    pretrain = False\n    batch_size = 2048\n    learning_rate = 1e-4\n    AdamW_decay = 1e-6\n    device = torch.device('cuda') if torch.cuda.is_available() else 'cpu'\n    folds = 10`\n\n`class OneDCNN(torch.nn.Module):\n    def __init__(self):\n        super().__init__()\n        \n        INPUT_DIM = 37\n        HIDDEN_DIM = 128\n        \n        self.embedding = torch.nn.Embedding(\n            num_embeddings=INPUT_DIM, \n            embedding_dim=HIDDEN_DIM,\n            padding_idx=0\n        )  # [BATCH_SIZE, 142] --> [BATCH_SIZE, 142, 128]\n        \n        self.model1 = torch.nn.Sequential(\n            torch.nn.Conv1d(128, 32, 3),  # [BATCH_SIZE, 128, 142] --> [BATCH_SIZE, 32, 140]\n            torch.nn.ReLU(),\n            torch.nn.Conv1d(32, 64, 3),  # [BATCH_SIZE, 32, 140] --> [BATCH_SIZE, 64, 138]\n            torch.nn.ReLU(),\n            torch.nn.Conv1d(64, 96, 3),  # [BATCH_SIZE, 64, 138] --> [BATCH_SIZE, 96, 136]\n            torch.nn.AdaptiveMaxPool1d(1)  # [BATCH_SIZE, 96, 136] --> [BATCH_SIZE, 96, 1]\n        )\n        \n        self.model2 = torch.nn.Sequential(\n            torch.nn.Linear(96, 1024),\n            torch.nn.ReLU(),\n            torch.nn.Linear(1024, 512),\n            torch.nn.ReLU(),\n            torch.nn.Linear(512, 3)\n        )\n    \n    def forward(self, x):\n        x = x.long()\n        x = self.embedding(x)  # [BATCH_SIZE, 142] --> [BATCH_SIZE, 142, 128]\n        x = x.permute(0, 2, 1)  # [BATCH_SIZE, 142, 128] --> [BATCH_SIZE, 128, 142]\n        x = self.model1(x)\n        x = x.squeeze(2)  # [BATCH_SIZE, 96, 1] --> [BATCH_SIZE, 96]\n        x = self.model2(x)\n        return x`\n\n`class Trainer:\n    def __init__(self, loaders, model):\n        self.train_loader, self.val_loader = loaders\n        self.model = model\n        self.optim = torch.optim.AdamW(self.model.parameters(), lr=CFG.learning_rate, weight_decay=CFG.AdamW_decay)\n        \n        self.train_loss = []\n        self.val_loss = []\n    \n    def clear(self):\n        gc.collect()\n        torch.cuda.empty_cache()\n    \n    def train_one_epoch(self):\n        running_loss = 0\n        progress = tqdm(self.train_loader, total=len(self.train_loader))\n        \n        for i, (features, targets) in enumerate(progress):\n            self.optim.zero_grad()\n            \n            targets_pred = self.model(features)\n            loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n            running_loss += loss.item()\n            \n            loss.backward()\n            self.optim.step()\n        \n        train_loss = running_loss / len(self.train_loader)\n        self.train_loss.append(train_loss)\n    \n    @torch.no_grad()\n    def val_one_epoch(self):\n        running_loss = 0\n        progress = tqdm(self.val_loader, total=len(self.val_loader))\n        \n        for (features, targets) in progress:\n            targets_pred = self.model(features)\n            loss = F.binary_cross_entropy_with_logits(targets_pred, targets)\n            running_loss += loss.item()\n            \n        val_loss = running_loss / len(self.val_loader)\n        self.val_loss.append(val_loss)\n    \n    def fit(self):\n        progress = tqdm(range(1, 2), desc='Training...')\n        \n        for epoch in progress:\n            progress.set_description(f\"EPOCH{epoch}/{1} | training...\")\n            self.train_one_epoch()\n            self.clear()\n        \n            progress.set_description(f\"EPOCH{epoch}/{1} | validating...\")\n            self.val_one_epoch()\n            self.clear()\n            \n            print(f\"{'-'*30} EPOCH {epoch} / {1} {'-'*30}\")\n            print(f\"train_loss: {self.train_loss[-1]}\")\n            print(f\"val_loss: {self.val_loss[-1]}\")\n            \n    def predict(self, test_loader):\n        preds = []\n        \n        for features in test_loader:\n            targets_pred = self.model(features)\n            preds.append(targets_pred.detach().cpu())\n        \n        preds = torch.concat(preds)\n        return preds\n    \n    def get_model(self):\n        return self.model`",
    "2898026": "Okay, I see you're new to the PyTorch ecosystem. Let me give you some pointers to help you out.\n\nFirst of all, data preprocessing is crucial. Make sure to scale or normalize your features using something like scikit-learn's StandardScaler. This can help your model converge faster and more stably.\n\nNext, your learning rate of 1e-4 might be a bit too high for such a large dataset. Try lowering it to 1e-5 or even 1e-6 and see if that improves things.\n\nThe large gap between your training and validation losses points to overfitting. Consider adding some regularization, like dropout layers or L2 regularization on the weights. Also, make sure your model's complexity is appropriate for your task - not too simple to underfit, but not too complex to overfit.\n\nCheck if your dataset has class imbalance. If some classes have way more samples than others, your model might be biased. You can fix this by oversampling the minority classes, undersampling the majority classes, or adjusting class weights in the loss function.\n\nMake sure you're using the right loss function. For multi-class classification, use CrossEntropyLoss instead of binary cross-entropy.\n\nTo make your life easier, use PyTorch Lightning. It's a lightweight PyTorch wrapper that handles a lot of the boring stuff for you. Here's how your model could look with Lightning:\n\n```python\nclass SimpleOneDCNN(LightningModule):\n    def __init__(self, input_dim=37, embedding_dim=128, num_classes=3):\n        super().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, 64, kernel_size=3) \n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d(1)\n        self.fc = nn.Linear(64, num_classes)\n        \n    def forward(self, x):\n        x = self.embedding(x) \n        x = torch.einsum(\"bse->bes\", x)  # transpose for Conv1d\n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze(2)\n        x = self.fc(x)\n        return x\n\n    def training_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('train_loss', loss)\n        return loss\n\n    def validation_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('val_loss', loss)\n\n    def configure_optimizers(self):\n        return torch.optim.AdamW(self.parameters(), lr=1e-5)\n```\n\nI snuck in an [https://pytorch.org/docs/stable/generated/torch.einsum.html](einsum) operation for the transpose in the forward method. Einsum is a super handy tool for tensor ops that can often be faster and cleaner than chaining together a bunch of separate PyTorch functions.\n\nHere's a minimum viable example:\n\n```python\nimport torch\nfrom torch import nn\nfrom torch.utils.data import DataLoader, TensorDataset\nfrom pytorch_lightning import LightningModule, Trainer\n\n\nclass SimpleOneDCNN(LightningModule):\n    def __init__(self, input_dim=37, embedding_dim=128, num_classes=3):\n        super().__init__()\n        self.embedding = nn.Embedding(input_dim, embedding_dim)\n        self.conv1 = nn.Conv1d(embedding_dim, 64, kernel_size=3)\n        self.relu = nn.ReLU()\n        self.pool = nn.AdaptiveMaxPool1d(1)\n        self.fc = nn.Linear(64, num_classes)\n\n    def forward(self, x):\n        x = self.embedding(x)\n        x = torch.einsum(\"bse->bes\", x)  # transpose for Conv1d\n        x = self.conv1(x)\n        x = self.relu(x)\n        x = self.pool(x)\n        x = x.squeeze(2)\n        x = self.fc(x)\n        return x\n\n    def training_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('train_loss', loss)\n        return loss\n\n    def validation_step(self, batch, batch_idx):\n        x, y = batch\n        y_hat = self(x)\n        loss = nn.CrossEntropyLoss()(y_hat, y)\n        self.log('val_loss', loss)\n\n    def configure_optimizers(self):\n        return torch.optim.AdamW(self.parameters(), lr=1e-3)\n\n\n# Example synthetic data\nX = torch.randint(0, 37, (1000, 142))\ny = torch.randint(0, 3, (1000,))\n\n# Create DataLoaders\ndataset = TensorDataset(X, y)\ntrain_loader = DataLoader(dataset, batch_size=32)\nval_loader = DataLoader(dataset, batch_size=32)\n\n# Train the model\nmodel = SimpleOneDCNN()\ntrainer = Trainer(max_epochs=5)\ntrainer.fit(model, train_loader, val_loader)\n```\n**One thing to note... don't focus too much on the actual *value* of the loss, just look for the loss to decrease in a steady way - watch for numerical stability above all else**\n\nGood luck :)",
    "2898499": "Thanks a lot ! I will do a quick study on Lightning immediately",
    "2899918": "Btw, feel free to take a look at my reimplementation of the 1dcnn notebook in PyTorch :)\n\n[Link](https://www.kaggle.com/code/yyyu54/pytorch-version-belka-1dcnn-starter-with-all-data)",
    "2900311": "YESYES !! I have tried the Pytorch Lightning Module, and just follow your wonderful notebook to design my new 1dcnn model. Now I can reach 0.386 on the public !! Thanks for your work :) Upvote !!",
    "2900318": "Actually, I have read your notebook a few days ago. But at that time, I was still unfamiliar with the Lightning Module. And now, I have found that it's much more convenient and powerful than the original one :) (By the way, you are 1st in this competition ! Omg. Amazing.",
    "2900551": "Glad it worked for you! Best of luck for the last couple of days!"
  },
  "source": "meta"
}