{
  "id": 163754,
  "title": "My special list of what can be done",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/163754",
  "author_name": "",
  "post_date": "2020-07-03T10:38:03.080533700Z",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all, I have to tell you that since I've started in this competition (12 days ago):\n1.  I have changed my ideas about  many important things\n2. I have learned a lot of things.\n3. I have lost many hours of sleep (but with fun)</p>\n\n<p>I want to try to summarize what I think are the most important things I have adopted so far:</p>\n\n<ul>\n<li>to be fast and make effective usage of TPU/GPU use <strong>TFRecords</strong>; At least with Tensorflow it is far by far too fast, no other way</li>\n<li>use <strong>TPU</strong>; I use GPU in my private environment only to test new ideas and code, then I go on Kaggle TPU</li>\n<li><strong>Augmented Dataset (30+30+10)</strong> in TFRecords format helps a lot in fighting over-fitting; Thanks to the guys who provided the datasets</li>\n<li>Use a double-head NN (metadata + image)</li>\n<li>use focal-loss to handle class imbalance</li>\n<li>take care of the learning rate and learning rate scheduling (I can still improve on this)</li>\n<li>take the BEST 4 models (255,384, 512, 768) and combine in a **super-classifier **using an ensemble</li>\n</ul>\n\n<p>Maybe as the next step, I should do CV, but it will make runs at least 3 times slower.</p>\n\n<p>Ok, for now it is 0.943 LB. Not bad, and I should be capable of almost easily improve two classifiers out of 4. </p>\n\n<p>What do you think?</p>",
  "messages": [
    {
      "id": "913653",
      "postDate": "07/03/2020 10:38:03",
      "content": "<p>First of all, I have to tell you that since I've started in this competition (12 days ago):\n1.  I have changed my ideas about  many important things\n2. I have learned a lot of things.\n3. I have lost many hours of sleep (but with fun)</p>\n\n<p>I want to try to summarize what I think are the most important things I have adopted so far:</p>\n\n<ul>\n<li>to be fast and make effective usage of TPU/GPU use <strong>TFRecords</strong>; At least with Tensorflow it is far by far too fast, no other way</li>\n<li>use <strong>TPU</strong>; I use GPU in my private environment only to test new ideas and code, then I go on Kaggle TPU</li>\n<li><strong>Augmented Dataset (30+30+10)</strong> in TFRecords format helps a lot in fighting over-fitting; Thanks to the guys who provided the datasets</li>\n<li>Use a double-head NN (metadata + image)</li>\n<li>use focal-loss to handle class imbalance</li>\n<li>take care of the learning rate and learning rate scheduling (I can still improve on this)</li>\n<li>take the BEST 4 models (255,384, 512, 768) and combine in a **super-classifier **using an ensemble</li>\n</ul>\n\n<p>Maybe as the next step, I should do CV, but it will make runs at least 3 times slower.</p>\n\n<p>Ok, for now it is 0.943 LB. Not bad, and I should be capable of almost easily improve two classifiers out of 4. </p>\n\n<p>What do you think?</p>",
      "rawMarkdown": "First of all, I have to tell you that since I've started in this competition (12 days ago):\n1.  I have changed my ideas about  many important things\n2. I have learned a lot of things.\n3. I have lost many hours of sleep (but with fun)\n\nI want to try to summarize what I think are the most important things I have adopted so far:\n\n- to be fast and make effective usage of TPU/GPU use **TFRecords**; At least with Tensorflow it is far by far too fast, no other way\n- use **TPU**; I use GPU in my private environment only to test new ideas and code, then I go on Kaggle TPU\n- **Augmented Dataset (30+30+10)** in TFRecords format helps a lot in fighting over-fitting; Thanks to the guys who provided the datasets\n- Use a double-head NN (metadata + image)\n- use focal-loss to handle class imbalance\n- take care of the learning rate and learning rate scheduling (I can still improve on this)\n- take the BEST 4 models (255,384, 512, 768) and combine in a **super-classifier **using an ensemble\n\nMaybe as the next step, I should do CV, but it will make runs at least 3 times slower.\n\nOk, for now it is 0.943 LB. Not bad, and I should be capable of almost easily improve two classifiers out of 4. \n\nWhat do you think?",
      "votes": null
    },
    {
      "id": "914003",
      "postDate": "07/03/2020 15:13:53",
      "content": "<p>Same results for me:\n-External Data\n-Focal Loss\n-Save model at different checkpoints to increase single model score\n-Use different size images at test time, kind of like TTA</p>\n\n<p>Right now im getting 0.930 lb single model single fold only with images, however im struggling to implement the meta data to my model. Im getting higher CV than my image model but the lb score is much lower (0.911).. The model also has a double NN head that combines into a classifier. Any tips would be appreciated! :)</p>",
      "rawMarkdown": "Same results for me:\n-External Data\n-Focal Loss\n-Save model at different checkpoints to increase single model score\n-Use different size images at test time, kind of like TTA\n\nRight now im getting 0.930 lb single model single fold only with images, however im struggling to implement the meta data to my model. Im getting higher CV than my image model but the lb score is much lower (0.911).. The model also has a double NN head that combines into a classifier. Any tips would be appreciated! :)",
      "votes": null
    },
    {
      "id": "914060",
      "postDate": "07/03/2020 15:47:21",
      "content": "<p>If you have different images from the same patient in both training and validation the meta-data may lead to over fitting the validation set (sex and age latently encode patient identity). As the patients between the training and test set are independent your model will not generalize as well and hence the lower LB. </p>\n\n<p>This could account for your problem, so if you haven't split your training/validation folds on patient_id it could be worth trying this to check. I can't imagine it will do any harm.</p>",
      "rawMarkdown": "If you have different images from the same patient in both training and validation the meta-data may lead to over fitting the validation set (sex and age latently encode patient identity). As the patients between the training and test set are independent your model will not generalize as well and hence the lower LB. \n\nThis could account for your problem, so if you haven't split your training/validation folds on patient_id it could be worth trying this to check. I can't imagine it will do any harm.",
      "votes": null
    },
    {
      "id": "914235",
      "postDate": "07/03/2020 17:33:24",
      "content": "<p>Unfortunately my data is split by patient_id but i should train on other folds to see the results </p>",
      "rawMarkdown": "Unfortunately my data is split by patient_id but i should train on other folds to see the results",
      "votes": null
    },
    {
      "id": "914282",
      "postDate": "07/03/2020 18:14:05",
      "content": "<p>At the beginning I had the same behavior, I thought it was useless to work with metadata. But then carefully looking at the data I realized that my NN was overfitting. I improved doing some hyperparameters tuning (dropout) and definitively with the ensemble.    </p>",
      "rawMarkdown": "At the beginning I had the same behavior, I thought it was useless to work with metadata. But then carefully looking at the data I realized that my NN was overfitting. I improved doing some hyperparameters tuning (dropout) and definitively with the ensemble.",
      "votes": null
    },
    {
      "id": "914376",
      "postDate": "07/03/2020 20:07:43",
      "content": "<p>Thanks for your advice! I will try to increase dropout probability and see if it changes my results :)</p>",
      "rawMarkdown": "Thanks for your advice! I will try to increase dropout probability and see if it changes my results :)",
      "votes": null
    },
    {
      "id": "914533",
      "postDate": "07/04/2020 02:23:07",
      "content": "<p>How do you save model checkpoints when using TPU?</p>",
      "rawMarkdown": "How do you save model checkpoints when using TPU?",
      "votes": null
    },
    {
      "id": "914580",
      "postDate": "07/04/2020 04:26:02",
      "content": "<p>Can you elaborate on saving checkpoints for a better single model score? Do you mean like save a model once every Epoch for example, and then try to score with each of them?</p>",
      "rawMarkdown": "Can you elaborate on saving checkpoints for a better single model score? Do you mean like save a model once every Epoch for example, and then try to score with each of them?",
      "votes": null
    },
    {
      "id": "914698",
      "postDate": "07/04/2020 06:42:16",
      "content": "<p>You can use code like this. It checks at the end of every epoch and keeps the two best models based on val_acc</p>\n\n<p>class save_best(tf.keras.callbacks.Callback):\n    def <strong>init</strong>(self, model):\n        self.model = model</p>\n\n<pre><code>def on_epoch_end(self, epoch, logs=None):\n\n    if (epoch &amp;gt; 0):\n        score = logs.get(\"val_auc\")\n    else:\n        score = -1\n\n    if (score &amp;gt; best_score.min()):\n\n        idx_min = np.argmin(best_score)\n\n        best_score[idx_min] = score\n        best_epoch[idx_min] = epoch + 1\n\n        path_best_model=f'best_model_{idx_min}.h5'\n        self.model.save(SAVE_DIR + \"/\" + path_best_model)\n</code></pre>\n\n<p>best_epoch = np.zeros(NBEST)\nbest_score = np.zeros(NBEST)</p>\n\n<p>hist = model.fit(get_training_dataset(train_dataset), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS,\n                 validation_data=get_validation_dataset(val_dataset),\n                 validation_steps=VALIDATION_STEPS_PER_EPOCH,\n                callbacks=[csv_logger, save_best(model), lr_scheduler])</p>",
      "rawMarkdown": "You can use code like this. It checks at the end of every epoch and keeps the two best models based on val_acc\n\nclass save_best(tf.keras.callbacks.Callback):\n    def __init__(self, model):\n        self.model = model\n\n    def on_epoch_end(self, epoch, logs=None):\n        \n        if (epoch &gt; 0):\n            score = logs.get(\"val_auc\")\n        else:\n            score = -1\n      \n        if (score &gt; best_score.min()):\n          \n            idx_min = np.argmin(best_score)\n\n            best_score[idx_min] = score\n            best_epoch[idx_min] = epoch + 1\n\n            path_best_model=f'best_model_{idx_min}.h5'\n            self.model.save(SAVE_DIR + \"/\" + path_best_model)\n\n\nbest_epoch = np.zeros(NBEST)\nbest_score = np.zeros(NBEST)\n\n\nhist = model.fit(get_training_dataset(train_dataset), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS,\n                 validation_data=get_validation_dataset(val_dataset),\n                 validation_steps=VALIDATION_STEPS_PER_EPOCH,\n                callbacks=[csv_logger, save_best(model), lr_scheduler])",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 914003,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "07/03/2020 15:13:53",
      "content": "<p>Same results for me:\n-External Data\n-Focal Loss\n-Save model at different checkpoints to increase single model score\n-Use different size images at test time, kind of like TTA</p>\n\n<p>Right now im getting 0.930 lb single model single fold only with images, however im struggling to implement the meta data to my model. Im getting higher CV than my image model but the lb score is much lower (0.911).. The model also has a double NN head that combines into a classifier. Any tips would be appreciated! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 914060,
          "author_name": "fchmiel",
          "author_url": "",
          "post_date": "07/03/2020 15:47:21",
          "content": "<p>If you have different images from the same patient in both training and validation the meta-data may lead to over fitting the validation set (sex and age latently encode patient identity). As the patients between the training and test set are independent your model will not generalize as well and hence the lower LB. </p>\n\n<p>This could account for your problem, so if you haven't split your training/validation folds on patient_id it could be worth trying this to check. I can't imagine it will do any harm.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914235,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "07/03/2020 17:33:24",
          "content": "<p>Unfortunately my data is split by patient_id but i should train on other folds to see the results </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914282,
          "author_name": "luigisaetta",
          "author_url": "",
          "post_date": "07/03/2020 18:14:05",
          "content": "<p>At the beginning I had the same behavior, I thought it was useless to work with metadata. But then carefully looking at the data I realized that my NN was overfitting. I improved doing some hyperparameters tuning (dropout) and definitively with the ensemble.    </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914376,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "07/03/2020 20:07:43",
          "content": "<p>Thanks for your advice! I will try to increase dropout probability and see if it changes my results :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914533,
          "author_name": "santiviquez",
          "author_url": "",
          "post_date": "07/04/2020 02:23:07",
          "content": "<p>How do you save model checkpoints when using TPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914580,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/04/2020 04:26:02",
          "content": "<p>Can you elaborate on saving checkpoints for a better single model score? Do you mean like save a model once every Epoch for example, and then try to score with each of them?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914698,
          "author_name": "luigisaetta",
          "author_url": "",
          "post_date": "07/04/2020 06:42:16",
          "content": "<p>You can use code like this. It checks at the end of every epoch and keeps the two best models based on val_acc</p>\n\n<p>class save_best(tf.keras.callbacks.Callback):\n    def <strong>init</strong>(self, model):\n        self.model = model</p>\n\n<pre><code>def on_epoch_end(self, epoch, logs=None):\n\n    if (epoch &amp;gt; 0):\n        score = logs.get(\"val_auc\")\n    else:\n        score = -1\n\n    if (score &amp;gt; best_score.min()):\n\n        idx_min = np.argmin(best_score)\n\n        best_score[idx_min] = score\n        best_epoch[idx_min] = epoch + 1\n\n        path_best_model=f'best_model_{idx_min}.h5'\n        self.model.save(SAVE_DIR + \"/\" + path_best_model)\n</code></pre>\n\n<p>best_epoch = np.zeros(NBEST)\nbest_score = np.zeros(NBEST)</p>\n\n<p>hist = model.fit(get_training_dataset(train_dataset), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS,\n                 validation_data=get_validation_dataset(val_dataset),\n                 validation_steps=VALIDATION_STEPS_PER_EPOCH,\n                callbacks=[csv_logger, save_best(model), lr_scheduler])</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "913653": "First of all, I have to tell you that since I've started in this competition (12 days ago):\n1.  I have changed my ideas about  many important things\n2. I have learned a lot of things.\n3. I have lost many hours of sleep (but with fun)\n\nI want to try to summarize what I think are the most important things I have adopted so far:\n\n- to be fast and make effective usage of TPU/GPU use **TFRecords**; At least with Tensorflow it is far by far too fast, no other way\n- use **TPU**; I use GPU in my private environment only to test new ideas and code, then I go on Kaggle TPU\n- **Augmented Dataset (30+30+10)** in TFRecords format helps a lot in fighting over-fitting; Thanks to the guys who provided the datasets\n- Use a double-head NN (metadata + image)\n- use focal-loss to handle class imbalance\n- take care of the learning rate and learning rate scheduling (I can still improve on this)\n- take the BEST 4 models (255,384, 512, 768) and combine in a **super-classifier **using an ensemble\n\nMaybe as the next step, I should do CV, but it will make runs at least 3 times slower.\n\nOk, for now it is 0.943 LB. Not bad, and I should be capable of almost easily improve two classifiers out of 4. \n\nWhat do you think?",
    "914003": "Same results for me:\n-External Data\n-Focal Loss\n-Save model at different checkpoints to increase single model score\n-Use different size images at test time, kind of like TTA\n\nRight now im getting 0.930 lb single model single fold only with images, however im struggling to implement the meta data to my model. Im getting higher CV than my image model but the lb score is much lower (0.911).. The model also has a double NN head that combines into a classifier. Any tips would be appreciated! :)",
    "914060": "If you have different images from the same patient in both training and validation the meta-data may lead to over fitting the validation set (sex and age latently encode patient identity). As the patients between the training and test set are independent your model will not generalize as well and hence the lower LB. \n\nThis could account for your problem, so if you haven't split your training/validation folds on patient_id it could be worth trying this to check. I can't imagine it will do any harm.",
    "914235": "Unfortunately my data is split by patient_id but i should train on other folds to see the results",
    "914282": "At the beginning I had the same behavior, I thought it was useless to work with metadata. But then carefully looking at the data I realized that my NN was overfitting. I improved doing some hyperparameters tuning (dropout) and definitively with the ensemble.",
    "914376": "Thanks for your advice! I will try to increase dropout probability and see if it changes my results :)",
    "914533": "How do you save model checkpoints when using TPU?",
    "914580": "Can you elaborate on saving checkpoints for a better single model score? Do you mean like save a model once every Epoch for example, and then try to score with each of them?",
    "914698": "You can use code like this. It checks at the end of every epoch and keeps the two best models based on val_acc\n\nclass save_best(tf.keras.callbacks.Callback):\n    def __init__(self, model):\n        self.model = model\n\n    def on_epoch_end(self, epoch, logs=None):\n        \n        if (epoch &gt; 0):\n            score = logs.get(\"val_auc\")\n        else:\n            score = -1\n      \n        if (score &gt; best_score.min()):\n          \n            idx_min = np.argmin(best_score)\n\n            best_score[idx_min] = score\n            best_epoch[idx_min] = epoch + 1\n\n            path_best_model=f'best_model_{idx_min}.h5'\n            self.model.save(SAVE_DIR + \"/\" + path_best_model)\n\n\nbest_epoch = np.zeros(NBEST)\nbest_score = np.zeros(NBEST)\n\n\nhist = model.fit(get_training_dataset(train_dataset), steps_per_epoch=STEPS_PER_EPOCH, epochs=EPOCHS,\n                 validation_data=get_validation_dataset(val_dataset),\n                 validation_steps=VALIDATION_STEPS_PER_EPOCH,\n                callbacks=[csv_logger, save_best(model), lr_scheduler])"
  },
  "source": "meta"
}