{"cells":[{"metadata":{"id":"2sK4N8oEpxrB","papermill":{"duration":0.023781,"end_time":"2020-12-29T00:23:17.836493","exception":false,"start_time":"2020-12-29T00:23:17.812712","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# Cassava RAPIDS kNN\n\n**The RAPIDS suite from NVIDIA is a collection of libraries that allows you to perform end-to-end data science pipelines entirely on GPUs. Leading Kaggle solutions and higher performing models/solutions rely on iterations; the more you can iterate and experiment, the more optimal solution you will find, so ensuring a fast pipeline is key. We can easily do this with RAPIDS libraries like `cuDF` and `cuML`. We will focus on the `cuML` library in this notebook, using `cuML KNN` for inference, and `cuML KMeans / T-SNE` to visualize how our data is clustered.**\n\n**We will extract image embeddings with EfficientNets and then use these CNN embeddings to train a RAPIDS cuML kNN to find similar images to see how kNN performs compared to more advanced models, like deep CNNs. This kernel is entirely motivated by [Chris Deotte](https://www.kaggle.com/cdeotte)'s notebook [here](https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates) from the Melanoma competition, in which he uses this technique to find duplicate images. Before giving this notebook an upvote (if you intend to), please give his one first.**\n\n**Note that I am using [Dimitre](https://www.kaggle.com/dimitreoliveira)'s TFRecords that can be found [here](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-tfrecords-512x512). He also has 128x128, 256x256, and 384x384 sized images that I added for experimental purposes. Please give his datasets an upvote (and his work in general, it is excellent).**"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:17.886860Z","iopub.status.busy":"2020-12-29T00:23:17.886210Z","iopub.status.idle":"2020-12-29T00:23:23.526903Z","shell.execute_reply":"2020-12-29T00:23:23.528056Z"},"papermill":{"duration":5.66903,"end_time":"2020-12-29T00:23:23.528259","exception":false,"start_time":"2020-12-29T00:23:17.859229","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"import sys\nsys.path.append('/kaggle/input/efficientnet-keras-dataset/efficientnet_kaggle')\nfrom efficientnet.tfkeras import *","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.execute_input":"2020-12-29T00:23:23.598689Z","iopub.status.busy":"2020-12-29T00:23:23.597898Z","iopub.status.idle":"2020-12-29T00:23:23.981862Z","shell.execute_reply":"2020-12-29T00:23:23.980592Z"},"id":"ot-sd-nvpxrB","outputId":"4cb746df-6d08-4e75-bb41-039877c2c2bb","papermill":{"duration":0.41736,"end_time":"2020-12-29T00:23:23.982002","exception":false,"start_time":"2020-12-29T00:23:23.564642","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"from matplotlib import pyplot as plt\nimport math, os, cv2, gc, re\nimport numpy as np, pandas as pd\nfrom time import time\nfrom sklearn.metrics import classification_report, accuracy_score\nfrom sklearn.model_selection import KFold\n\nimport tensorflow as tf\nimport tensorflow.keras.backend as K","execution_count":null,"outputs":[]},{"metadata":{"id":"OoU1g5QwpxrF","papermill":{"duration":0.023193,"end_time":"2020-12-29T00:23:24.028017","exception":false,"start_time":"2020-12-29T00:23:24.004824","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# I. Configuration"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:24.089757Z","iopub.status.busy":"2020-12-29T00:23:24.088878Z","iopub.status.idle":"2020-12-29T00:23:24.938841Z","shell.execute_reply":"2020-12-29T00:23:24.938151Z"},"id":"EkGzSMH4pxrF","outputId":"dd30b419-a0b7-41d5-8e6c-2e30824b1511","papermill":{"duration":0.887665,"end_time":"2020-12-29T00:23:24.938966","exception":false,"start_time":"2020-12-29T00:23:24.051301","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"DEVICE = 'GPU'\n\nif DEVICE == \"TPU\":\n    print(\"connecting to TPU...\")\n    try:\n        tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n        print('Running on TPU ', tpu.master())\n    except ValueError:\n        print(\"Could not connect to TPU\")\n        tpu = None\n\n    if tpu:\n        try:\n            print(\"initializing  TPU ...\")\n            tf.config.experimental_connect_to_cluster(tpu)\n            tf.tpu.experimental.initialize_tpu_system(tpu)\n            strategy = tf.distribute.experimental.TPUStrategy(tpu)\n            print(\"TPU initialized\")\n        except _:\n            print(\"failed to initialize TPU\")\n    else:\n        DEVICE = \"GPU\"\n\nif DEVICE != \"TPU\":\n    print(\"Using default strategy for CPU and single GPU\")\n    strategy = tf.distribute.get_strategy()\n\nif DEVICE == \"GPU\":\n    print(\"Num GPUs Available: \", len(tf.config.experimental.list_physical_devices('GPU')))\n    \n\nAUTO = tf.data.experimental.AUTOTUNE\nREPLICAS = strategy.num_replicas_in_sync\nprint(f'REPLICAS: {REPLICAS}')","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:24.990120Z","iopub.status.busy":"2020-12-29T00:23:24.989446Z","iopub.status.idle":"2020-12-29T00:23:24.994154Z","shell.execute_reply":"2020-12-29T00:23:24.993603Z"},"id":"xpYJbPsFpxrM","papermill":{"duration":0.031705,"end_time":"2020-12-29T00:23:24.994256","exception":false,"start_time":"2020-12-29T00:23:24.962551","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"SEED = 34 \n             \nIMAGE_SIZE = [256, 256]               \n\nBATCH_SIZE = 16 * REPLICAS \n\nFOLDS = 5\n\nVERBOSE = 1","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.023117,"end_time":"2020-12-29T00:23:25.040663","exception":false,"start_time":"2020-12-29T00:23:25.017546","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# II. Dataset Functions\n\n**Note that we are using `noisy-student` pretrained weights, so we need to normalize our inputs by subtracting ImageNet mean `.449`and dividing by standard deviation `.226`. We must do this because the EfficientNet is not being trained anymore: if it was, the models weights would correct themselves. You can experiment with training on Cassava data before extracting embeddings: you will get different features and perhaps they will do better than just ImageNet pretraining.**"},{"metadata":{"_kg_hide-input":false,"execution":{"iopub.execute_input":"2020-12-29T00:23:25.106725Z","iopub.status.busy":"2020-12-29T00:23:25.105762Z","iopub.status.idle":"2020-12-29T00:23:25.119416Z","shell.execute_reply":"2020-12-29T00:23:25.119898Z"},"id":"7PEHqJZ0pxrR","papermill":{"duration":0.056217,"end_time":"2020-12-29T00:23:25.120020","exception":false,"start_time":"2020-12-29T00:23:25.063803","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"def decode_image(image_data):\n    image = tf.image.decode_jpeg(image_data, channels=3)\n    image = ((tf.cast(image, tf.float32) / 255.0) - 0.449) / 0.226     \n    #image = tf.cast(image, tf.float32) / 255.0\n    image = tf.reshape(image, [*IMAGE_SIZE, 3])\n    return image\n\ndef read_labeled_tfrecord(example, return_image_name):\n    LABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string), \n        \"target\": tf.io.FixedLenFeature([], tf.int64), \n        \"image_name\": tf.io.FixedLenFeature([], tf.string)\n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    label = tf.cast(example['target'], tf.int32)\n    img_name = example['image_name']\n    \n    if return_image_name:\n        return image, label, img_name\n    else:\n        return image, label\n\ndef read_unlabeled_tfrecord(example, return_image_name):\n    UNLABELED_TFREC_FORMAT = {\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"id\": tf.io.FixedLenFeature([], tf.string), \n    }\n    \n    example = tf.io.parse_single_example(example, UNLABELED_TFREC_FORMAT)\n    image = decode_image(example['image'])\n    idnum = example['id']\n    return image, idnum if return_image_name else 0\n\ndef get_dataset(files, shuffle=False, repeat=False, labeled=True, return_image_names=True,\n                batch_size=BATCH_SIZE, dim=IMAGE_SIZE[0]):\n   \n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=AUTO)\n\n    if repeat:\n        ds = ds.repeat()\n    \n    if shuffle: \n        ds = ds.shuffle(2048)\n        opt = tf.data.Options()\n        opt.experimental_deterministic = False\n        ds = ds.with_options(opt)\n        \n    if labeled: \n        ds = ds.map(lambda example: read_labeled_tfrecord(example, return_image_names), \n                    num_parallel_calls=AUTO)  \n    else:\n        ds = ds.map(lambda example: read_unlabeled_tfrecord(example, return_image_names), \n                    num_parallel_calls=AUTO)  \n\n    ds = ds.batch(batch_size)\n    ds = ds.prefetch(AUTO)\n    \n    return ds\n\ndef load_image(jpeg_path, image_id):  \n    img = ((cv2.imread(os.path.join(jpeg_path, image_id))/255.0) - .449) / .226\n    #img = cv2.imread(os.path.join(jpeg_path, image_id))/255.0\n    img = cv2.resize(img, (IMAGE_SIZE[0], IMAGE_SIZE[1]))[:, :, ::-1]\n    return img\n\ndef generator(filepath, paths, batch_size=BATCH_SIZE):\n    i=0\n    while i <= len(paths):\n        batch = []\n        for cpt in range(batch_size):\n            if i + cpt >= len(paths):\n                i += batch_size\n                break\n            batch.append(load_image(filepath, paths[i+cpt]))\n            \n        i += batch_size\n        yield np.stack(batch)\n\ndef count_data_items(filenames):\n    n = [int(re.compile(r'-([0-9]*)\\.').search(filename).group(1)) for filename in filenames]\n    return np.sum(n)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:25.169077Z","iopub.status.busy":"2020-12-29T00:23:25.168427Z","iopub.status.idle":"2020-12-29T00:23:25.199689Z","shell.execute_reply":"2020-12-29T00:23:25.199061Z"},"papermill":{"duration":0.057915,"end_time":"2020-12-29T00:23:25.199793","exception":false,"start_time":"2020-12-29T00:23:25.141878","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"TRAINING_FILENAMES =  tf.io.gfile.glob(f'../input/cassava-leaf-disease-tfrecords-{IMAGE_SIZE[0]}x{IMAGE_SIZE[0]}' + '/*.tfrec')\nTRAINING_FILENAMES_ORG = tf.io.gfile.glob('../input/cassava-leaf-disease-classification/train_tfrecords' + '/*.tfrec')\nTEST_JPEG_PATH = \"../input/cassava-leaf-disease-classification/test_images\"\n\nsubmission = pd.read_csv('../input/cassava-leaf-disease-classification/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":0.021949,"end_time":"2020-12-29T00:23:25.244731","exception":false,"start_time":"2020-12-29T00:23:25.222782","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# III. RAPIDS cuML kNN\n\n**Now we can extract image embeddings with an EfficientNet and use RAPIDS cuML kNN for quick GPU model training. The kNN training doesn't take long at all, it is the feature extraction that takes a while. You can save the embeddings and upload them later to make the below loop faster.**"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:25.295451Z","iopub.status.busy":"2020-12-29T00:23:25.294425Z","iopub.status.idle":"2020-12-29T00:23:25.297522Z","shell.execute_reply":"2020-12-29T00:23:25.296993Z"},"papermill":{"duration":0.030216,"end_time":"2020-12-29T00:23:25.297618","exception":false,"start_time":"2020-12-29T00:23:25.267402","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"FINE_TUNE = True\nFT_EPOCHS = 3","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:25.347731Z","iopub.status.busy":"2020-12-29T00:23:25.347117Z","iopub.status.idle":"2020-12-29T00:23:28.794298Z","shell.execute_reply":"2020-12-29T00:23:28.793783Z"},"papermill":{"duration":3.473533,"end_time":"2020-12-29T00:23:28.794483","exception":false,"start_time":"2020-12-29T00:23:25.320950","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"import cuml\nprint('RAPIDS version',cuml.__version__)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:28.862298Z","iopub.status.busy":"2020-12-29T00:23:28.856350Z","iopub.status.idle":"2020-12-29T00:23:28.864662Z","shell.execute_reply":"2020-12-29T00:23:28.865258Z"},"papermill":{"duration":0.047269,"end_time":"2020-12-29T00:23:28.865404","exception":false,"start_time":"2020-12-29T00:23:28.818135","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"def efficientnet(b, image_size, head=False, LR=5e-4):\n    efns = [EfficientNetB0, EfficientNetB1, EfficientNetB2,\n            EfficientNetB3, EfficientNetB4, EfficientNetB5,\n            EfficientNetB6]\n    with strategy.scope():\n        efficient = efns[b](\n            input_shape=(image_size, image_size, 3),\n            weights='noisy-student', #imagenet\n            include_top=False\n        )\n        efficient.trainable=True\n        \n        if head:\n            model = tf.keras.Sequential([\n                efficient,\n                tf.keras.layers.GlobalAveragePooling2D(name='pooling'), \n                tf.keras.layers.Dropout(.2), \n                tf.keras.layers.Dense(5, activation='softmax')\n            ])\n            \n        else:\n            model = tf.keras.Sequential([\n                efficient,\n                tf.keras.layers.GlobalAveragePooling2D()]) \n                \n    if head: model.compile(optimizer=tf.keras.optimizers.Adam(LR), \n                           loss='sparse_categorical_crossentropy',\n                           metrics=['sparse_categorical_accuracy'])\n        \n    return model","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:28.914644Z","iopub.status.busy":"2020-12-29T00:23:28.913914Z","iopub.status.idle":"2020-12-29T00:23:28.917048Z","shell.execute_reply":"2020-12-29T00:23:28.916563Z"},"papermill":{"duration":0.029519,"end_time":"2020-12-29T00:23:28.917152","exception":false,"start_time":"2020-12-29T00:23:28.887633","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"FOLDS = 5\nSEED = 34\nN_NEIGH = 10","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T00:23:28.992493Z","iopub.status.busy":"2020-12-29T00:23:28.991426Z","iopub.status.idle":"2020-12-29T01:47:44.093834Z","shell.execute_reply":"2020-12-29T01:47:44.093245Z"},"papermill":{"duration":5055.154631,"end_time":"2020-12-29T01:47:44.093960","exception":false,"start_time":"2020-12-29T00:23:28.939329","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"skf = KFold(n_splits=FOLDS,shuffle=True,random_state=SEED)\npreds_all = []\npreds_model = []\noof_pred = []\noof_labels = []\n\nfor f, (train_index, val_index) in enumerate(skf.split(TRAINING_FILENAMES)):\n    \n    print('#'*30); print('#### FOLD',f+1); print('#'*30); print('')\n    print('Getting datasets...'); print('')\n    \n    train_ds = get_dataset(list(pd.DataFrame({'TRAINING_FILENAMES': TRAINING_FILENAMES}).loc[train_index]['TRAINING_FILENAMES']),\n                                            labeled=True, return_image_names=False, repeat=False, shuffle=False)\n    val_ds = get_dataset(list(pd.DataFrame({'TRAINING_FILENAMES': TRAINING_FILENAMES}).loc[val_index]['TRAINING_FILENAMES']),\n                                            labeled=True, return_image_names=False, repeat=False, shuffle=False)\n    test_ds = generator(TEST_JPEG_PATH, submission.image_id.values)\n    train_labs = [target.numpy() for img, target in iter(train_ds.unbatch())]\n    val_labs = [target.numpy() for img, target in iter(val_ds.unbatch())]\n    \n    effnet_ = efficientnet(b=3, image_size=IMAGE_SIZE[0], head=FINE_TUNE)\n    \n    if FINE_TUNE:\n        train_ds_ = get_dataset(list(pd.DataFrame({'TRAINING_FILENAMES': TRAINING_FILENAMES}).loc[train_index]['TRAINING_FILENAMES']),\n                                            labeled=True, return_image_names=False, repeat=True, shuffle=True)\n        \n        print('Fine tuning EfficientNet...'); print('')\n        ct_train = count_data_items(list(pd.DataFrame({'TRAINING_FILENAMES': TRAINING_FILENAMES}).loc[train_index]['TRAINING_FILENAMES']))\n        effnet_.fit(train_ds_, \n                    validation_data=val_ds,\n                    verbose=1, \n                    steps_per_epoch=ct_train//BATCH_SIZE,\n                    epochs=FT_EPOCHS)\n        print('')\n        effnet = tf.keras.Model(inputs = effnet_.input, \n               outputs = effnet_.get_layer('pooling').output)\n\n    else: effnet = effnet_\n\n    print('Getting embeddings...'); print('')\n    embed = effnet.predict(train_ds, verbose=1)\n    embed_val = effnet.predict(val_ds, verbose=1)\n    embed_test = effnet.predict(test_ds, verbose=1)\n    np.save(f'embed_b4_{f}_{IMAGE_SIZE[0]}',embed.astype('float32'))\n    np.save(f'embed_val_b4_{f}_{IMAGE_SIZE[0]}',embed_val.astype('float32'))\n    \n    print(''); print('Training and inferring...'); print('')\n    model = cuml.neighbors.KNeighborsClassifier(n_neighbors=N_NEIGH)\n    model.fit(embed, np.array(train_labs))\n    print('Training and inference complete.'); print('')\n    \n    preds = model.predict_proba(embed_test)\n    preds_model.append(preds) \n    \n    acc = accuracy_score(model.predict(embed_val), np.array(val_labs))\n    print(f'Fold {f + 1} accuracy: {acc}'); print('')\n    \n    oof_labels.append([target.numpy() for img, target in iter(val_ds.unbatch())])\n    x_oof = val_ds.map(lambda image, image_name: image)\n    oof_pred.append(model.predict(embed_val))\n    \npreds_model = np.stack(preds_model).mean(0)\npreds_all.append(preds_model)\npreds_all = np.stack(preds_all)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:47:57.994260Z","iopub.status.busy":"2020-12-29T01:47:57.993357Z","iopub.status.idle":"2020-12-29T01:47:58.080291Z","shell.execute_reply":"2020-12-29T01:47:58.072572Z"},"papermill":{"duration":7.264795,"end_time":"2020-12-29T01:47:58.080490","exception":false,"start_time":"2020-12-29T01:47:50.815695","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"y_true = np.concatenate(oof_labels)\ny_preds = np.concatenate(oof_pred)\n\nprint(classification_report(y_true, y_preds))\nprint(f\"OOF accuracy score: {accuracy_score(y_true, y_preds)}\")","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":7.374252,"end_time":"2020-12-29T01:48:13.007220","exception":false,"start_time":"2020-12-29T01:48:05.632968","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# IV. RAPIDS cuML KMeans\n\n**The embeddings generated by the EfficientNetB4 are dimension `1792`, so each image is represented as a point in `1792` dimensional space. We can cluster these points together with a RAPIDS KMeans model and then view what the images in each cluster look like. For now, we will force `N_CLUSTERS == 5`.**"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:48:26.885630Z","iopub.status.busy":"2020-12-29T01:48:26.884612Z","iopub.status.idle":"2020-12-29T01:50:37.184895Z","shell.execute_reply":"2020-12-29T01:50:37.125121Z"},"papermill":{"duration":137.06304,"end_time":"2020-12-29T01:50:37.185033","exception":false,"start_time":"2020-12-29T01:48:20.121993","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train_dummy = get_dataset(TRAINING_FILENAMES, labeled=True,\n                         return_image_names=True, repeat=False, \n                         shuffle=False)\nnames = np.array([img_name.numpy().decode(\"utf-8\") for img, label, img_name in iter(train_dummy.unbatch())])\nlabels = np.array([label.numpy() for img, label, img_name in iter(train_dummy.unbatch())])\ntrain_full = get_dataset(TRAINING_FILENAMES, labeled=True,\n                       return_image_names=False, repeat=False, \n                       shuffle=False)\nembed_full = effnet.predict(train_full, verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:50:52.092947Z","iopub.status.busy":"2020-12-29T01:50:52.092302Z","iopub.status.idle":"2020-12-29T01:50:52.100109Z","shell.execute_reply":"2020-12-29T01:50:52.100557Z"},"papermill":{"duration":7.62451,"end_time":"2020-12-29T01:50:52.100696","exception":false,"start_time":"2020-12-29T01:50:44.476186","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"train = pd.DataFrame()\ntrain['image_id'] = names\ntrain['label'] = labels","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:51:07.107767Z","iopub.status.busy":"2020-12-29T01:51:07.106957Z","iopub.status.idle":"2020-12-29T01:51:07.477470Z","shell.execute_reply":"2020-12-29T01:51:07.477962Z"},"papermill":{"duration":8.154264,"end_time":"2020-12-29T01:51:07.478092","exception":false,"start_time":"2020-12-29T01:50:59.323828","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"N_CLUSTERS = 5\nmodel = cuml.KMeans(n_clusters=N_CLUSTERS)\nmodel.fit(embed_full)\ntrain['cluster'] = model.labels_\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:51:22.147639Z","iopub.status.busy":"2020-12-29T01:51:22.146800Z","iopub.status.idle":"2020-12-29T01:51:27.215049Z","shell.execute_reply":"2020-12-29T01:51:27.215742Z"},"papermill":{"duration":12.442475,"end_time":"2020-12-29T01:51:27.215895","exception":false,"start_time":"2020-12-29T01:51:14.773420","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"JPEG_TRAIN = '../input/cassava-leaf-disease-classification/train_images/'\n\nfor k in range(N_CLUSTERS):\n    print('#'*25);\n    print(f'#### Cluster {k} of similar train images')\n    print('#'*25)\n    df = train.loc[train.cluster==k]\n    plt.figure(figsize=(20,10))\n    for j in range(8):\n        plt.subplot(2,4,j+1)\n        img = cv2.imread(JPEG_TRAIN+names[df.index[j]])\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n        plt.axis('off')\n        plt.title(f\"{names[df.index[j]]}, Target = {df.loc[df.index[j],'label']}\")\n        plt.imshow(img)  \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":7.013949,"end_time":"2020-12-29T01:51:41.494470","exception":false,"start_time":"2020-12-29T01:51:34.480521","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# V. RAPIDS cuML T-SNE\n\n**We can of course also project these points in `1792` dimensional space to a 2 dimensional space to see how are samples are clustered. Forgive the ensuing mess of colors...**"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:51:56.569194Z","iopub.status.busy":"2020-12-29T01:51:56.566971Z","iopub.status.idle":"2020-12-29T01:51:56.570076Z","shell.execute_reply":"2020-12-29T01:51:56.570645Z"},"papermill":{"duration":7.819576,"end_time":"2020-12-29T01:51:56.570797","exception":false,"start_time":"2020-12-29T01:51:48.751221","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"PERPLEXITY = 5","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:52:11.920018Z","iopub.status.busy":"2020-12-29T01:52:11.918636Z","iopub.status.idle":"2020-12-29T01:52:14.395456Z","shell.execute_reply":"2020-12-29T01:52:14.394902Z"},"papermill":{"duration":9.851361,"end_time":"2020-12-29T01:52:14.395566","exception":false,"start_time":"2020-12-29T01:52:04.544205","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"model = cuml.TSNE(perplexity=PERPLEXITY)\nembed2D = model.fit_transform(embed_full)\ntrain['x'] = embed2D[:,0]\ntrain['y'] = embed2D[:,1]","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:52:29.328686Z","iopub.status.busy":"2020-12-29T01:52:29.327661Z","iopub.status.idle":"2020-12-29T01:52:29.767294Z","shell.execute_reply":"2020-12-29T01:52:29.767802Z"},"papermill":{"duration":8.099249,"end_time":"2020-12-29T01:52:29.767934","exception":false,"start_time":"2020-12-29T01:52:21.668685","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(10,10))\ndf1 = train.loc[train.label==0]\ndf2 = train.loc[train.label==1]\ndf3 = train.loc[train.label==2]\ndf4 = train.loc[train.label==3]\ndf5 = train.loc[train.label==4]\n\nplt.scatter(df1.x,df1.y,c='blue',s=10,label='0')\nplt.scatter(df2.x,df2.y,c='cornflowerblue',s=10,label='1')\nplt.scatter(df3.x,df3.y,c='purple',s=10,label='2')\nplt.scatter(df4.x,df4.y,c='steelblue',s=10,label='3')\nplt.scatter(df5.x,df5.y,c='mediumpurple',s=10,label='4')\nplt.legend();","execution_count":null,"outputs":[]},{"metadata":{"papermill":{"duration":7.292661,"end_time":"2020-12-29T01:52:44.144408","exception":false,"start_time":"2020-12-29T01:52:36.851747","status":"completed"},"tags":[]},"cell_type":"markdown","source":"# VI. Submission"},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:52:59.263854Z","iopub.status.busy":"2020-12-29T01:52:59.261189Z","iopub.status.idle":"2020-12-29T01:52:59.267840Z","shell.execute_reply":"2020-12-29T01:52:59.266265Z"},"papermill":{"duration":7.827483,"end_time":"2020-12-29T01:52:59.267939","exception":false,"start_time":"2020-12-29T01:52:51.440456","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"print(preds_all.shape)\npreds_all","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:53:14.339150Z","iopub.status.busy":"2020-12-29T01:53:14.338314Z","iopub.status.idle":"2020-12-29T01:53:14.519531Z","shell.execute_reply":"2020-12-29T01:53:14.518656Z"},"papermill":{"duration":7.538288,"end_time":"2020-12-29T01:53:14.519651","exception":false,"start_time":"2020-12-29T01:53:06.981363","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"submission[\"label\"] = preds_all.mean(0).argmax(1)\nsubmission.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"execution":{"iopub.execute_input":"2020-12-29T01:53:28.942255Z","iopub.status.busy":"2020-12-29T01:53:28.941627Z","iopub.status.idle":"2020-12-29T01:53:28.947149Z","shell.execute_reply":"2020-12-29T01:53:28.946683Z"},"papermill":{"duration":7.286377,"end_time":"2020-12-29T01:53:28.947281","exception":false,"start_time":"2020-12-29T01:53:21.660904","status":"completed"},"tags":[],"trusted":false},"cell_type":"code","source":"submission","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}