{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🧰 Imports and Setups","metadata":{}},{"cell_type":"markdown","source":"## 🔵 1. Install Weights and Biases\n\nWeights & Biases comes baked into your Kaggle kernels! Since W&B is rapidly improving pip installing the latest version is is recommended. ","metadata":{}},{"cell_type":"code","source":"!pip install wandb","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 🔵 2. Import wandb and login\n\nAn unique API token is required to login to Weights and Biases. \n\n1. If you don't have a Weights and Biases account you can head over to https://wandb.ai/site and create a FREE account.\n2. In order to access your API token you can head over to https://wandb.ai/authorize.\n\nThere are two ways you can login using a Kaggle kernel:\n\n1. Run a cell with `wandb.login()`. It will ask for the API key to login.\n2. You can use Kaggle secrets to store your API key and use the code snippet below to login. Check out this [discussion post](https://www.kaggle.com/product-feedback/114053) to learn more about Kaggle secrets. \n\n```\nfrom kaggle_secrets import UserSecretsClient\n\nuser_secrets = UserSecretsClient()\n\n# I have saved my API token with \"wandb_api\" as Label. \n# If you use some other Label make sure to change the same below. \nwandb_api = user_secrets.get_secret(\"wandb_api\") \n\nwandb.login(key=wandb_api)\n```\nMore on W&B login [here](https://docs.wandb.ai/ref/cli/wandb-login).","metadata":{}},{"cell_type":"code","source":"import wandb\nfrom wandb.keras import WandbCallback\n\n\n\nwandb.login()\n# use this key=de280857522249143800325c5025e1fe19c0e6ae\n# then you will see result there https://wandb.ai/joannawozna/plant-pathology","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 📌 Pro tip: If you don't indend to make your kernel public use Kaggle secrets to login. For a public facing kernel use `wandb.login()`.\n\n> 📌 Note: `WandbCallback()` is a light weight Keras callback baked with goodness of Weights and Biases. We will not focus on this callback in this kernel. However, if you are interested to learn more check out the [official docs](https://docs.wandb.ai/guides/integrations/keras). ","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nprint(tf.__version__)\nfrom tensorflow.keras import layers\nfrom tensorflow.keras import models\nimport tensorflow_addons as tfa\n\nimport os\nimport json\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nfrom sklearn.model_selection import train_test_split","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set the random seeds\ndef seed_everything():\n    os.environ['TF_CUDNN_DETERMINISTIC'] = '1' \n    np.random.seed(hash(\"improves reproducibility\") % 2**32 - 1)\n    tf.random.set_seed(hash(\"by removing stochasticity\") % 2**32 - 1)\n    \nseed_everything()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 📀 Hyperparameters","metadata":{}},{"cell_type":"code","source":"TRAIN_PATH = '../input/resized-plant2021/img_sz_256/'\nAUTOTUNE = tf.data.experimental.AUTOTUNE\n\nCONFIG = dict (\n    num_labels = 6,\n    train_val_split = 0.2,\n    img_width = 224,\n    img_height = 224,\n    batch_size = 64,\n    epochs = 20,\n    learning_rate = 0.001,\n    architecture = \"CNN\",\n    infra = \"Kaggle\",\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 📌 Pro tip: Use a dictionary structure to store your hyperparameters. You can easily log your config dict to W&B. It's a good practice to save your training configuration to analyze your experiments and reproduce your work in future. ","metadata":{}},{"cell_type":"markdown","source":"# 🔨 Build Input Pipeline","metadata":{}},{"cell_type":"code","source":"label_to_id = {\n    'healthy': 0,\n    'scab': 1,\n    'frog_eye_leaf_spot': 2,\n    'rust': 3,\n    'complex': 4,\n    'powdery_mildew': 5\n}\n\nid_to_label = {value:key for key, value in label_to_id.items()} \n\ndef make_path(row):\n    return TRAIN_PATH+row.image\n\ndef parse_labels(row):\n    label_list = row.labels.split()\n    labels = []\n    for label in label_list:\n        labels.append(str(label_to_id[label]))\n    \n    return ' '.join(labels)\n\n# Read train.csv file\ndf = pd.read_csv('../input/plant-pathology-2021-fgvc8/train.csv')\n# Get absolute path\ndf['image'] = df.apply(lambda row: make_path(row), axis=1)\n# Parse labels\ndf['labels'] = df.apply(lambda row: parse_labels(row), axis=1)\n\n# Look at the dataframe\ndf.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df, valid_df = train_test_split(df, test_size=CONFIG['train_val_split'])\nprint(f'Number of train images: {len(train_df)} and validation images: {len(valid_df)}')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"@tf.function\ndef decode_image(image):\n    # convert the compressed string to a 3D uint8 tensor\n    image = tf.image.decode_jpeg(image, channels=3)\n    # Normalize image\n    image = tf.image.convert_image_dtype(image, dtype=tf.float32)\n    # resize the image to the desired size\n    return image\n\n@tf.function\ndef load_image(df_dict):\n    # Load image\n    image = tf.io.read_file(df_dict['image'])\n    image = decode_image(image)\n    \n    # Resize image\n    image = tf.image.resize(image, (CONFIG['img_height'], CONFIG['img_width']))\n    \n    # Parse label\n    label = tf.strings.split(df_dict['labels'], sep='')\n    label = tf.strings.to_number(label, out_type=tf.int32)\n    label = tf.reduce_sum(tf.one_hot(indices=label, depth=CONFIG['num_labels']), axis=0)\n    \n    return image, label","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"AUTOTUNE = tf.data.AUTOTUNE\n\ntrainloader = tf.data.Dataset.from_tensor_slices(dict(train_df))\nvalidloader = tf.data.Dataset.from_tensor_slices(dict(valid_df))\n\ntrainloader = (\n    trainloader\n    .shuffle(1024)\n    .map(load_image, num_parallel_calls=AUTOTUNE)\n    .batch(CONFIG['batch_size'])\n    .prefetch(AUTOTUNE)\n)\n\nvalidloader = (\n    validloader\n    .map(load_image, num_parallel_calls=AUTOTUNE)\n    .batch(CONFIG['batch_size'])\n    .prefetch(AUTOTUNE)\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_batch(image_batch, label_batch):\n    plt.figure(figsize=(20,20))\n    for n in range(25):\n        ax = plt.subplot(5,5,n+1)\n        plt.imshow(image_batch[n])\n        plt.title(' '.join([id_to_label[i] for i, label in enumerate(label_batch[n].numpy()) if label==1.]))\n        plt.axis('off')\n\nimage_batch, label_batch = next(iter(trainloader))\nshow_batch(image_batch, label_batch)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 🐤 Model","metadata":{}},{"cell_type":"code","source":"def get_model():\n    base_model = tf.keras.applications.EfficientNetB0(include_top=False, weights='imagenet')\n    base_model.trainabe = True\n\n    inputs = layers.Input((CONFIG['img_height'], CONFIG['img_width'], 3))\n    x = base_model(inputs, training=True)\n    x = layers.GlobalAveragePooling2D()(x)\n    x = layers.Dropout(0.5)(x)\n    outputs = layers.Dense(len(label_to_id), activation='sigmoid')(x)\n    \n    return models.Model(inputs, outputs)\n\ntf.keras.backend.clear_session()\nmodel = get_model()\nmodel.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#### EfficientNetB1\ndef get_model():\n    base_model = tf.keras.applications.EfficientNetB1(include_top=False, weights='imagenet')\n    base_model.trainabe = True\n\n    inputs = layers.Input((CONFIG['img_height'], CONFIG['img_width'], 3))\n    x = base_model(inputs, training=True)\n    x = layers.GlobalAveragePooling2D()(x)\n    x = layers.Dropout(0.5)(x)\n    outputs = layers.Dense(len(label_to_id), activation='sigmoid')(x)\n    \n    return models.Model(inputs, outputs)\n\ntf.keras.backend.clear_session()\nmodel = get_model()\nmodel.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" ","metadata":{}},{"cell_type":"markdown","source":"# 🚄 Train with Weights and Biases\n\nInstrumenting Weights and Biases in your training script is very easy.  the simple building blocks to track an experiment with W&B.\n\n* `wandb.init()`: Initialize a new run at the top of your script. This returns a Run object and creates a local directory where all logs and files are saved, then streamed asynchronously to a W&B server. [Check out docs](https://docs.wandb.ai/guides/track/launch).\n\n* `wandb.config`: Save a dictionary of hyperparameters such as learning rate or model type. The model settings you capture in config are useful later to organize and query your results. [Check out docs](https://docs.wandb.ai/guides/track/config).\n\n* `wandb.log()`: Log metrics over time in a training loop, such as accuracy and loss. [Check out docs](https://docs.wandb.ai/guides/track/log).\n\n* `wandb.log_artifact`: Save outputs of a run, like the model weights or a table of predictions. This lets you track not just model training, but all the pipeline steps that affect the final model. [Check out docs](https://docs.wandb.ai/guides/artifacts).\n\n","metadata":{}},{"cell_type":"code","source":"# Initialize model\ntf.keras.backend.clear_session()\nmodel = get_model()\n\n# Compile model\noptimizer = tf.keras.optimizers.Adam(learning_rate=CONFIG['learning_rate'])\nmodel.compile(optimizer, \n              loss=tfa.losses.SigmoidFocalCrossEntropy(), \n              metrics=[tf.keras.metrics.AUC(multi_label=True), tfa.metrics.F1Score(num_classes=6, average='micro')])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 🔵 3. Use `wandb.init()` to initialize a new W&B run.\n\nIn an ML training pipeline, you could add `wandb.init()` to the beginning of your training script as well as your evaluation script, and each piece would be tracked as a run in W&B.\n\nTake a note of these arguments. \n\n* `entity`: An entity is a username or team name where you're sending runs. \n* `project`: The name of the project where you're sending the new run. If the project is not specified, the run is put in an \"Uncategorized\" project.\n* `config`: This sets wandb.config, a dictionary-like object for saving inputs to your job, like hyperparameters for a model or settings for a data preprocessing job. \n* `group`: Specify a group to organize individual runs into a larger experiment.This is a super handy feature. For example, you can create group for different model architecture names. \n* `job_type`: Specify the type of run, which is useful when you're grouping runs together into larger experiments using group. Typical job types are \"train\", \"evaluate\", etc. ","metadata":{}},{"cell_type":"code","source":"# Update CONFIG dict with the name of the model.\nCONFIG['model_name'] = 'efficientnetb0'\nprint('Training configuration: ', CONFIG)\n\n# Initialize W&B run\nrun = wandb.init(project='plant-pathology', \n                 config=CONFIG,\n                 group='EfficientNetB0', \n                 job_type='train')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 📌 Note: If you don't provide `entity` W&B will use the default entity. If you have multiple W&B user/team account you can visit https://wandb.ai/settings to set the default entity.\n\n> 📌 Pro tip: You can update your config even after `wandb.init()` is called. ","metadata":{}},{"cell_type":"markdown","source":"## 🔵 4. Use `wandb.config` to update your logged CONFIG.\n\nSet the wandb.config object in your script to save your training config: hyperparameters, input settings like dataset name or model type, and any other independent variables for your experiments. This is useful for analyzing your experiments and reproducing your work in the future. You'll be able to group by config values in the web interface, comparing the settings of different runs and seeing how these affect the output. Check out [the docs to learn more](https://docs.wandb.ai/guides/track/config). \n","metadata":{}},{"cell_type":"code","source":"wandb.config.type = 'baseline'\nwandb.config.kaggle_competition = 'Plant Pathology 2021 - FGVC8'","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 📌 Pro tip: There are multiple ways to update your config. Using config with Weights and Biases has a lot of advantages. Check out the [docs](https://docs.wandb.ai/guides/track/config) and this [colab notebook](https://colab.research.google.com/github/wandb/examples/blob/master/colabs/wandb-log/Configs_in_W%26B.ipynb#scrollTo=xFf3zjBSixC1) to learn more. ","metadata":{}},{"cell_type":"markdown","source":"![img](https://i.imgur.com/WxLjIBx.gif)","metadata":{}},{"cell_type":"markdown","source":"Weights and Biases comes with a light weight integration for Keras. We will be using W&B Keras integration (`WandbCallback()`) to automatically save all the metrics and the loss values tracked in `model.fit()`. Check out the [docs to learn more about this integration](https://docs.wandb.ai/guides/integrations/keras). ","metadata":{}},{"cell_type":"code","source":"earlystopper = tf.keras.callbacks.EarlyStopping(\n    monitor='val_loss', patience=10, verbose=0, mode='min',\n    restore_best_weights=True\n)\n\n# Train\nmodel.fit(trainloader, \n          epochs=CONFIG['epochs'],\n          validation_data=validloader,\n          callbacks=[WandbCallback(),\n                     earlystopper])\n\n# Close W&B run\nmodel.save(\"./\" + str(CONFIG['model_name']))\nrun.finish()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 📌 Pro tip: Head over to the W&B dashboard my clicking on the link generated above. \n\n> 📌 Pro tip: If you want to silence W&B related logs use this code snippet `os.environ[WANDB_SILENT] = \"true\"` after `import os`. Check out this [Stackoverflow answer](https://stackoverflow.com/a/65997094/8663152) for more details. \n\n> 📌 Pro tip: Use `run.finish()` to close the initialized W&B run after a `job_type` is finished. ","metadata":{}},{"cell_type":"markdown","source":"![img](https://i.imgur.com/Eq8X9RN.gif)","metadata":{}},{"cell_type":"markdown","source":"## 🔵 5. Using `wandb.log()` to log evaluation score.\n\nCall `wandb.log(dict)` to log a dictionary of metrics or custom objects to a step. Each time we log, the step is incremented by default, letting us view metrics over time. We can log (almost) anything - image, audio, video, segmentation masks, bounding boxes, HTML, 3D cloud points, molecules, etc. Check out the [official doc to see what all can we achieve (log) using W&B](https://docs.wandb.ai/guides/track/log). Alternatively check out this [YouTube video on logging (almost) anything with W&B](https://www.youtube.com/watch?v=96MxRvx15Ts). ","metadata":{}},{"cell_type":"code","source":"reconstructed_model = tf.keras.models.load_model(\"./\" + str(CONFIG['model_name']))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Initialize a new W&B run\nrun = wandb.init(project='plant-pathology', \n                 config=CONFIG,\n                 group='EfficientNetB0', \n                 job_type='evaluate') # Note the job_type\n\n# Configuration\nwandb.config.type = 'baseline'\nwandb.config.kaggle_competition = 'Plant Pathology 2021 - FGVC8'\n\n# Evaluate model\nloss, auc, f1_score = reconstructed_model.evaluate(validloader)\n\n# Log scores using wandb.log()\nwandb.log({'val_AUC': auc, \n           'val_F1_score': f1_score})\n\n# Close W&B run\nrun.finish()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![img](https://i.imgur.com/OBL9F1i.gif)\n\n> 📌 Pro tip: Note that the bar chart apprears if there are more than one value for a key. ","metadata":{}},{"cell_type":"markdown","source":"# 💾 Save Model Weights\n\n## 🔵 5. Use `wandb.log_artifacts()` to save your hard work.\n\nUse W&B Artifacts for dataset versioning, model versioning, and tracking dependencies and results across machine learning pipelines. Think of an artifact as a versioned folder of data. You can store entire datasets directly in artifacts, or use artifact references to point to data in other systems like S3, GCP, or your own system. A typical machine learning pipeline can be represented using Weights and Biases Data+Model Versioning tool called Artifacts. Learn more about [artifacts here](https://docs.wandb.ai/guides/artifacts).\n\n![img](https://i.imgur.com/dhntZxK.png)","metadata":{}},{"cell_type":"code","source":"# Save model\nmodel.save('efficientnetb0-baseline.h5')\n\n# Initialize a new W&B run\nrun = wandb.init(project='plant-pathology', \n                 config=CONFIG,\n                 group='EfficientNetB0', \n                 job_type='save') # Note the job_type\n\n# Configuration\nwandb.config.type = 'baseline'\nwandb.config.kaggle_competition = 'Plant Pathology 2021 - FGVC8'\n\n# Save as Model artifact\nartifact = wandb.Artifact('efficientnet-b0', type='model')\nartifact.add_file('efficientnetb0-baseline.h5')\nrun.log_artifact(artifact)\n\n# Close W&B run\nrun.finish()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![img](https://i.imgur.com/HjHDoSx.gif)","metadata":{}},{"cell_type":"markdown","source":"## 🔵 6. Filtering and Grouping\n\nYou can use filter and group feature on W&B dashboard to either hide crashed run, group together multiple runs under one experiment, group as per the type of the run, select runs that satisfy a condition, etc. You can learn more about Group feature [here](https://docs.wandb.ai/guides/track/advanced/grouping).\n\n![img](https://i.imgur.com/BeYKbfS.gif)","metadata":{}},{"cell_type":"markdown","source":"# ❄️ Resources\n\nI hope you find this kernel useful and will encouage you to try out Weights and Biases. Here are some relevant links that you might want to check out:\n\n* Check out the [official documentation](https://docs.wandb.ai/) to learn more about the best practices and advanced features. \n\n* Check out the [examples GitHub repository](https://github.com/wandb/examples) for curated and minimal examples. This can be a good starting point. \n\n* [Weights and Biases Fully Connected](https://wandb.ai/fully-connected) is a home for curated tutorials, free-form dicussions, paper summaries, industry expert advices and more. \n\nHere are some other Kaggle kernels instrumented with Weights and Biases that you might find useful. \n\n* [EfficientNet+Mixup+K-Fold using TF and wandb](https://www.kaggle.com/ayuraj/efficientnet-mixup-k-fold-using-tf-and-wandb)\n\n* [HPA: Segmentation Mask Visualization with W&B](https://www.kaggle.com/ayuraj/hpa-segmentation-mask-visualization-with-w-b)\n\n* [HPA: Multi-Label Classification with TF and W&B](https://www.kaggle.com/ayuraj/hpa-multi-label-classification-with-tf-and-w-b)\n\n* [🐦BirdCLEF: Quick EDA with W&B](https://www.kaggle.com/ayuraj/birdclef-quick-eda-with-w-b)","metadata":{}}]}