{
  "id": 261595,
  "title": "Any success with non-pretrained models?",
  "url": "/competitions/seti-breakthrough-listen/discussion/261595",
  "author_name": "Markus Frank",
  "post_date": "2021-08-04T17:26:40.421000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I was experimenting with 256x256 images of the signal parts only (AAA). A pretrained EfficientNetB1 achieves a 5-fold CV AUC of ~0.80 on this dataset without any augmentations. However, I haven't managed to get <em>any</em> non-pretrained model to learn anything at all. No matter if it's a simple 3-layer CNN with a MLP on top or a randomly initialized EfficientNet: it may memorize the train set if I remove all Dropouts, but validation AUC stays flat at 0.5. I don't expect top results, but no learning at all?</p>\n<p>Anybody seen this too? Is this to be expected somehow?</p>",
  "messages": [
    {
      "id": 1448386,
      "postDate": "2021-08-04T17:26:40.420Z",
      "content": "<p>I was experimenting with 256x256 images of the signal parts only (AAA). A pretrained EfficientNetB1 achieves a 5-fold CV AUC of ~0.80 on this dataset without any augmentations. However, I haven't managed to get <em>any</em> non-pretrained model to learn anything at all. No matter if it's a simple 3-layer CNN with a MLP on top or a randomly initialized EfficientNet: it may memorize the train set if I remove all Dropouts, but validation AUC stays flat at 0.5. I don't expect top results, but no learning at all?</p>\n<p>Anybody seen this too? Is this to be expected somehow?</p>",
      "rawMarkdown": "I was experimenting with 256x256 images of the signal parts only (AAA). A pretrained EfficientNetB1 achieves a 5-fold CV AUC of ~0.80 on this dataset without any augmentations. However, I haven't managed to get *any* non-pretrained model to learn anything at all. No matter if it's a simple 3-layer CNN with a MLP on top or a randomly initialized EfficientNet: it may memorize the train set if I remove all Dropouts, but validation AUC stays flat at 0.5. I don't expect top results, but no learning at all?\n\nAnybody seen this too? Is this to be expected somehow?",
      "votes": 6
    },
    {
      "id": 1448443,
      "postDate": "2021-08-04T17:36:07.817Z",
      "content": "<p>It takes some time, much longer then pretrained. Accuracy in my experiments was also worse. Dataset is specific, so this is kinda expected</p>",
      "rawMarkdown": "It takes some time, much longer then pretrained. Accuracy in my experiments was also worse. Dataset is specific, so this is kinda expected",
      "votes": 2,
      "replies": [
        {
          "id": 1448484,
          "postDate": "2021-08-04T17:41:23.543Z",
          "content": "<p>I'm training on TPU, so I tried for 100+ epochs once, just to check. While val loss decreases somewhat, val AUC oscillates around 0.5. Have you managed to achieve any success with a non-pretrained model? It seems that the features of the signal cannot be learned from the dataset at all but must come pretrained in the model.</p>",
          "rawMarkdown": "I'm training on TPU, so I tried for 100+ epochs once, just to check. While val loss decreases somewhat, val AUC oscillates around 0.5. Have you managed to achieve any success with a non-pretrained model? It seems that the features of the signal cannot be learned from the dataset at all but must come pretrained in the model."
        },
        {
          "id": 1448525,
          "postDate": "2021-08-04T17:50:37.663Z",
          "content": "<p>~.80 CV with light models (resnet18d, effnetb0) without pretrain, yes. val AUC stays at .5 for a long time, but eventually starts learning.  Cant really say specific params like epoch number, accuracy was too low to continue experiments</p>",
          "rawMarkdown": "~.80 CV with light models (resnet18d, effnetb0) without pretrain, yes. val AUC stays at .5 for a long time, but eventually starts learning.  Cant really say specific params like epoch number, accuracy was too low to continue experiments",
          "votes": 2
        },
        {
          "id": 1454848,
          "postDate": "2021-08-06T10:54:06.443Z",
          "content": "<p>You're right. My learning rate was way too low. Now it's finally learning. Interesting loss profile:<br>\n<img src=\"https://i.imgur.com/24FsVGO.png\" alt=\"img\"><br>\nSimple CNN. No augmentation. x axis is epochs, left is loss (blue train, orange val), right is auc.<br>\nYou really need to wait… ;)</p>",
          "rawMarkdown": "You're right. My learning rate was way too low. Now it's finally learning. Interesting loss profile:\n![img](https://i.imgur.com/24FsVGO.png)\nSimple CNN. No augmentation. x axis is epochs, left is loss (blue train, orange val), right is auc.\nYou really need to wait... ;)",
          "votes": 3
        },
        {
          "id": 1454927,
          "postDate": "2021-08-06T11:40:10.433Z",
          "content": "<p>Nice, doing lr scheduler can help to get basic understanding of what loss landscape is. For example fastai guys implemented \"lr-finder\", one epoch lr scheduler between extreme values, which supposed to find somewhat \"optimal\"  starting lr. But still moving  distributions of all nn parameter takes time.</p>",
          "rawMarkdown": "Nice, doing lr scheduler can help to get basic understanding of what loss landscape is. For example fastai guys implemented \"lr-finder\", one epoch lr scheduler between extreme values, which supposed to find somewhat \"optimal\"  starting lr. But still moving  distributions of all nn parameter takes time."
        }
      ]
    },
    {
      "id": 1451080,
      "postDate": "2021-08-05T08:55:56.703Z",
      "content": "<p>Not sure if it will work here, but maybe Self Training with Noisy Student ? <br>\n<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a><br>\nIt was used in the 6th place for RANZCR - in write-up \"Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle.\"</p>",
      "rawMarkdown": "Not sure if it will work here, but maybe Self Training with Noisy Student ? \nhttps://arxiv.org/pdf/1911.04252.pdf\nIt was used in the 6th place for RANZCR - in write-up \"Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle.\"",
      "votes": 1,
      "replies": [
        {
          "id": 1458520,
          "postDate": "2021-08-07T23:28:30.500Z",
          "content": "<p>Thank you for sharing paper. </p>\n<p>I also share the link of the discussion of 6th place solution of RANZCR:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616</a></p>\n<p>In this competition, I saw many train image that contains no distinguishable signal by human eye, so I think the training set is somewhat noisy.</p>",
          "rawMarkdown": "Thank you for sharing paper. \n\nI also share the link of the discussion of 6th place solution of RANZCR:\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616\n\nIn this competition, I saw many train image that contains no distinguishable signal by human eye, so I think the training set is somewhat noisy."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1448443,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2021-08-04T17:36:07.817000",
      "content": "<p>It takes some time, much longer then pretrained. Accuracy in my experiments was also worse. Dataset is specific, so this is kinda expected</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1448484,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-04T17:41:23.543000",
          "content": "<p>I'm training on TPU, so I tried for 100+ epochs once, just to check. While val loss decreases somewhat, val AUC oscillates around 0.5. Have you managed to achieve any success with a non-pretrained model? It seems that the features of the signal cannot be learned from the dataset at all but must come pretrained in the model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1448525,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-04T17:50:37.663000",
          "content": "<p>~.80 CV with light models (resnet18d, effnetb0) without pretrain, yes. val AUC stays at .5 for a long time, but eventually starts learning.  Cant really say specific params like epoch number, accuracy was too low to continue experiments</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1454848,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-06T10:54:06.443000",
          "content": "<p>You're right. My learning rate was way too low. Now it's finally learning. Interesting loss profile:<br>\n<img src=\"https://i.imgur.com/24FsVGO.png\" alt=\"img\"><br>\nSimple CNN. No augmentation. x axis is epochs, left is loss (blue train, orange val), right is auc.<br>\nYou really need to wait… ;)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1454927,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-06T11:40:10.433000",
          "content": "<p>Nice, doing lr scheduler can help to get basic understanding of what loss landscape is. For example fastai guys implemented \"lr-finder\", one epoch lr scheduler between extreme values, which supposed to find somewhat \"optimal\"  starting lr. But still moving  distributions of all nn parameter takes time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1451080,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-08-05T08:55:56.703000",
      "content": "<p>Not sure if it will work here, but maybe Self Training with Noisy Student ? <br>\n<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a><br>\nIt was used in the 6th place for RANZCR - in write-up \"Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle.\"</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1458520,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2021-08-07T23:28:30.500000",
          "content": "<p>Thank you for sharing paper. </p>\n<p>I also share the link of the discussion of 6th place solution of RANZCR:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/226616</a></p>\n<p>In this competition, I saw many train image that contains no distinguishable signal by human eye, so I think the training set is somewhat noisy.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1448386": "I was experimenting with 256x256 images of the signal parts only (AAA). A pretrained EfficientNetB1 achieves a 5-fold CV AUC of ~0.80 on this dataset without any augmentations. However, I haven't managed to get *any* non-pretrained model to learn anything at all. No matter if it's a simple 3-layer CNN with a MLP on top or a randomly initialized EfficientNet: it may memorize the train set if I remove all Dropouts, but validation AUC stays flat at 0.5. I don't expect top results, but no learning at all?\n\nAnybody seen this too? Is this to be expected somehow?",
    "1448443": "It takes some time, much longer then pretrained. Accuracy in my experiments was also worse. Dataset is specific, so this is kinda expected",
    "1451080": "Not sure if it will work here, but maybe Self Training with Noisy Student ? \nhttps://arxiv.org/pdf/1911.04252.pdf\nIt was used in the 6th place for RANZCR - in write-up \"Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle.\""
  }
}