{
  "id": 163395,
  "title": "how to split images for training between two categories?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/163395",
  "author_name": "",
  "post_date": "2020-07-01T21:07:02.184863700Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am new here and my question is how i can split images between melanoma and normal skin provided JPEG folder?</p>",
  "messages": [
    {
      "id": "911592",
      "postDate": "07/01/2020 21:07:02",
      "content": "<p>I am new here and my question is how i can split images between melanoma and normal skin provided JPEG folder?</p>",
      "rawMarkdown": "I am new here and my question is how i can split images between melanoma and normal skin provided JPEG folder?",
      "votes": null
    },
    {
      "id": "911670",
      "postDate": "07/01/2020 23:27:57",
      "content": "<p>You have to use the csv file given with the labels. The way I did it was read all of the rows of the csv file and make 2 arrays that hold the image names and the corresponding labels. Then loop through all the jpg files and get just the image name without the \".jpg\" extension. Using Numpy you can find the index of that name in your image name array and then find the corresponding label in the label array.</p>",
      "rawMarkdown": "You have to use the csv file given with the labels. The way I did it was read all of the rows of the csv file and make 2 arrays that hold the image names and the corresponding labels. Then loop through all the jpg files and get just the image name without the \".jpg\" extension. Using Numpy you can find the index of that name in your image name array and then find the corresponding label in the label array.",
      "votes": null
    },
    {
      "id": "914673",
      "postDate": "07/04/2020 06:17:32",
      "content": "<p>thank you so much man! i have another question that accuracy shoots to 0.98 and validation to 0.96 in first epoch. is there something wrong?</p>",
      "rawMarkdown": "thank you so much man! i have another question that accuracy shoots to 0.98 and validation to 0.96 in first epoch. is there something wrong?",
      "votes": null
    },
    {
      "id": "914683",
      "postDate": "07/04/2020 06:25:59",
      "content": "<p>If you use TFRecords then each image has its label included within the TFRecord. Kaggle datasets are at the following links <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256</a> </p>",
      "rawMarkdown": "If you use TFRecords then each image has its label included within the TFRecord. Kaggle datasets are at the following links [768x768][4], [512x512][1], [384x384][3], [256x256][2] \n\n[1]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[2]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[3]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[4]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[5]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images",
      "votes": null
    },
    {
      "id": "914769",
      "postDate": "07/04/2020 08:00:54",
      "content": "<p>Thank you so much ! i am an undergraduate student and just start learning this thing. do you make your own architectures for this type of competitions?  if so i am eager to learn how to design own architecture. can you help me with some study materiel? </p>",
      "rawMarkdown": "Thank you so much ! i am an undergraduate student and just start learning this thing. do you make your own architectures for this type of competitions?  if so i am eager to learn how to design own architecture. can you help me with some study materiel?",
      "votes": null
    },
    {
      "id": "915223",
      "postDate": "07/04/2020 15:08:36",
      "content": "<p>When doing image classification tasks where the number of training data is less than 100,000 images, it is best to use transfer learning. You download an architecture from the internet and it comes pretrained on 10,000,000 images!</p>\n\n<p>Then you build your own <strong>head</strong> architecture (a new top on the network) and finetune it on the competition data. For example, in TensorFlow, you start with</p>\n\n<pre><code>!pip install efficientnet\nimport efficientnet.tfkeras as efn\ninp = tf.keras.Input(shape=(1024,1024,3))\nbase_model = efn.EfficientNetB4(weights='imagenet',include_top=False) \nx = base_model(inp)\nx = tf.keras.layers.GlobalAveragePooling2D()(x) \n</code></pre>\n\n<p>Then you build your own architecture to finish the CNN. For example, the simpliest finish is just</p>\n\n<pre><code>x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\nmodel = tf.keras.Model(inputs=inp,outputs=x)\n</code></pre>",
      "rawMarkdown": "When doing image classification tasks where the number of training data is less than 100,000 images, it is best to use transfer learning. You download an architecture from the internet and it comes pretrained on 10,000,000 images!\n\nThen you build your own **head** architecture (a new top on the network) and finetune it on the competition data. For example, in TensorFlow, you start with\n\n    !pip install efficientnet\n    import efficientnet.tfkeras as efn\n    inp = tf.keras.Input(shape=(1024,1024,3))\n    base_model = efn.EfficientNetB4(weights='imagenet',include_top=False) \n    x = base_model(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x) \n\nThen you build your own architecture to finish the CNN. For example, the simpliest finish is just\n\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)",
      "votes": null
    },
    {
      "id": "915225",
      "postDate": "07/04/2020 15:09:43",
      "content": "<p>Review the public notebooks in this competition. They have examples in TensorFlow, PyTorch and examples with GPU and TPU. Next, a strategy for experimentation and increasing your LB score is shown in Bengali Comp <a href=\"https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop\">here</a></p>",
      "rawMarkdown": "Review the public notebooks in this competition. They have examples in TensorFlow, PyTorch and examples with GPU and TPU. Next, a strategy for experimentation and increasing your LB score is shown in Bengali Comp [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 911670,
      "author_name": "hmorera",
      "author_url": "",
      "post_date": "07/01/2020 23:27:57",
      "content": "<p>You have to use the csv file given with the labels. The way I did it was read all of the rows of the csv file and make 2 arrays that hold the image names and the corresponding labels. Then loop through all the jpg files and get just the image name without the \".jpg\" extension. Using Numpy you can find the index of that name in your image name array and then find the corresponding label in the label array.</p>",
      "votes": null,
      "replies": [
        {
          "id": 914673,
          "author_name": "mtalhaarshad",
          "author_url": "",
          "post_date": "07/04/2020 06:17:32",
          "content": "<p>thank you so much man! i have another question that accuracy shoots to 0.98 and validation to 0.96 in first epoch. is there something wrong?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 914683,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/04/2020 06:25:59",
      "content": "<p>If you use TFRecords then each image has its label included within the TFRecord. Kaggle datasets are at the following links <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">768x768</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">512x512</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">384x384</a>, <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">256x256</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 914769,
          "author_name": "mtalhaarshad",
          "author_url": "",
          "post_date": "07/04/2020 08:00:54",
          "content": "<p>Thank you so much ! i am an undergraduate student and just start learning this thing. do you make your own architectures for this type of competitions?  if so i am eager to learn how to design own architecture. can you help me with some study materiel? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915223,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/04/2020 15:08:36",
          "content": "<p>When doing image classification tasks where the number of training data is less than 100,000 images, it is best to use transfer learning. You download an architecture from the internet and it comes pretrained on 10,000,000 images!</p>\n\n<p>Then you build your own <strong>head</strong> architecture (a new top on the network) and finetune it on the competition data. For example, in TensorFlow, you start with</p>\n\n<pre><code>!pip install efficientnet\nimport efficientnet.tfkeras as efn\ninp = tf.keras.Input(shape=(1024,1024,3))\nbase_model = efn.EfficientNetB4(weights='imagenet',include_top=False) \nx = base_model(inp)\nx = tf.keras.layers.GlobalAveragePooling2D()(x) \n</code></pre>\n\n<p>Then you build your own architecture to finish the CNN. For example, the simpliest finish is just</p>\n\n<pre><code>x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\nmodel = tf.keras.Model(inputs=inp,outputs=x)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915225,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/04/2020 15:09:43",
          "content": "<p>Review the public notebooks in this competition. They have examples in TensorFlow, PyTorch and examples with GPU and TPU. Next, a strategy for experimentation and increasing your LB score is shown in Bengali Comp <a href=\"https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "911592": "I am new here and my question is how i can split images between melanoma and normal skin provided JPEG folder?",
    "911670": "You have to use the csv file given with the labels. The way I did it was read all of the rows of the csv file and make 2 arrays that hold the image names and the corresponding labels. Then loop through all the jpg files and get just the image name without the \".jpg\" extension. Using Numpy you can find the index of that name in your image name array and then find the corresponding label in the label array.",
    "914673": "thank you so much man! i have another question that accuracy shoots to 0.98 and validation to 0.96 in first epoch. is there something wrong?",
    "914683": "If you use TFRecords then each image has its label included within the TFRecord. Kaggle datasets are at the following links [768x768][4], [512x512][1], [384x384][3], [256x256][2] \n\n[1]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[2]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[3]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[4]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[5]: https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images",
    "914769": "Thank you so much ! i am an undergraduate student and just start learning this thing. do you make your own architectures for this type of competitions?  if so i am eager to learn how to design own architecture. can you help me with some study materiel?",
    "915223": "When doing image classification tasks where the number of training data is less than 100,000 images, it is best to use transfer learning. You download an architecture from the internet and it comes pretrained on 10,000,000 images!\n\nThen you build your own **head** architecture (a new top on the network) and finetune it on the competition data. For example, in TensorFlow, you start with\n\n    !pip install efficientnet\n    import efficientnet.tfkeras as efn\n    inp = tf.keras.Input(shape=(1024,1024,3))\n    base_model = efn.EfficientNetB4(weights='imagenet',include_top=False) \n    x = base_model(inp)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x) \n\nThen you build your own architecture to finish the CNN. For example, the simpliest finish is just\n\n    x = tf.keras.layers.Dense(1, activation='sigmoid')(x)\n    model = tf.keras.Model(inputs=inp,outputs=x)",
    "915225": "Review the public notebooks in this competition. They have examples in TensorFlow, PyTorch and examples with GPU and TPU. Next, a strategy for experimentation and increasing your LB score is shown in Bengali Comp [here][1]\n\n[1]: https://www.kaggle.com/cdeotte/how-to-compete-with-gpus-workshop"
  },
  "source": "meta"
}