{"cells":[{"metadata":{},"cell_type":"markdown","source":"## 参考\nhttps://www.kaggle.com/tanulsingh077/how-to-become-leaf-doctor-with-deep-learning"},{"metadata":{},"cell_type":"markdown","source":"# このコンペティションについて \n\nメラノーマに続いて、今年もまた古典的なコンピュータビジョンの分類問題が出題されました。CV を始めたばかりの人にとっては、このライブコンペに手を出してみて、最初の一歩を踏み出すことができる絶好の機会です。このコンテストでは、分類の精度が問われますが、それがどのくらいの頻度で起こるのでしょうか？\n\n通常、実際に必要なのは、葉の医者になることであり、農家が感染性の葉を識別し、手頃なレートでそれらを治すのを助けることです 😛 。\n \n# このノートブックについて\n\n* いつものように、これは初心者向けのノートで、キャッサバの葉の病気に特化した葉の医者に効率的になれる方法をお伝えします 😛 そして、主な方法論はディープラーニングです。\n\n* 私はあなたが知っておく必要があるすべてのものをカバーします , 専門知識から方法論まで , 私は問題を解決するために提案するさまざまなアイデアのベースラインの例と一緒に\n\n* 多くの混乱がなければ、あなたはこのノートブックに従うと、これはあなたの最初のCV competitionを行うことができます。\n\n* 機械学習とkaggleが全く初めての方は、私が書いたこの[guide](https://www.kaggle.com/tanulsingh077/tackling-any-kaggle-competition-the-noob-s-way) を見てみてください。 \n\n# Step 1 : 患者の分析\n\n* 彼のクライアントが話すことができないという事実を考慮して何かの前に葉の医者がすべきであることを最初のステップは何でしょうか？答えは当然簡単で、患者を見て何が間違っているかを分析します。\n\n* しかし、医師はそれを見るだけで何かが間違っているかどうかをどのように理解しているのでしょうか？医師としてこのためには、彼は正常な患者/葉がどのように見えるかを知っている必要がありますし、感染したものから健康な患者を分離するために、通常の動作からの逸脱（パターン、色、質感など）を観察する必要があります。今、さらに病気の特定のクラスに感染したものを分類するために、医師はまた、患者/葉の状態が異なる疾患のように見える方法を知っておく必要があります。\n\nこれらのポイントを念頭に置いて、基本的な馴染みのあるものから始めてみましょう。"},{"metadata":{"trusted":true},"cell_type":"code","source":"import sys\n#ライブラリをインポートする際にパスを追加\nsys.path.append('../input/pytorch-image-models/pytorch-image-models-master')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# Preliminaries\nimport os\nfrom pathlib import Path\nimport glob\n#プログレスバーを表示する\nfrom tqdm import tqdm\ntqdm.pandas()\nimport json\nimport pandas as pd\nimport numpy as np\n\n## Image hash\nimport imagehash\n\n# Visuals and CV2\nimport seaborn as sn\nimport matplotlib.pyplot as plt\nimport cv2\nfrom PIL import Image\n\n\n# albumentations for augs\n# 画像の加工\nimport albumentations\nfrom albumentations.pytorch.transforms import ToTensorV2\n\n# クラスタリング、次元削減\nfrom sklearn.cluster import KMeans\nfrom sklearn.decomposition import PCA\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import accuracy_score\n\n# Keras and TensorFlow\nfrom keras.preprocessing.image import load_img\nfrom keras.preprocessing.image import img_to_array \nfrom keras.applications.resnet50 import preprocess_input \n\n# models \nfrom keras.applications.resnet50 import ResNet50\nfrom keras.models import Model\n\n#torch\nimport torch\nimport timm\nimport torch\nimport torch.nn as nn\nfrom torch.nn import functional as F\nfrom torch.utils.data import Dataset,DataLoader","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Utils\n\nユーティリティー機能の項"},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_images(class_id, label, images_number,verbose=0):\n    '''\n    Courtesy of https://www.kaggle.com/isaienkov/cassava-leaf-disease-classification-data-analysis\n    '''\n    '''\n    ラベルがclass_idの画像をimages_number枚ランダムで取得し表示\n    '''\n    plot_list = train[train[\"label\"] == class_id].sample(images_number)['image_id'].tolist()\n    \n    # 画像のリストを表示\n    if verbose:\n        print(plot_list)\n        \n    labels = [label for i in range(len(plot_list))]\n    size = np.sqrt(images_number)\n    #画像をsubplotを使って複数表示するために、sizeという変数をうまく設定している\n    if int(size)*int(size) < images_number:\n        size = int(size) + 1\n        \n    plt.figure(figsize=(20, 20))\n    \n    for ind, (image_id, label) in enumerate(zip(plot_list, labels)):\n        plt.subplot(size, size, ind + 1)\n        image = cv2.imread(str(BASE_DIR/'train_images'/image_id))\n        #OpenCVでは画像をBGRの順で読み込むので変換する必要がある。\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n\n        plt.imshow(image)\n        plt.title(label, fontsize=12)\n        plt.axis(\"off\")\n    \n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"BASE_DIR = Path('../input/cassava-leaf-disease-classification')\n\n## Reading DataFrame having Labels\ntrain = pd.read_csv(BASE_DIR/'train.csv')\n\n## Label Mappings\nwith open(BASE_DIR/'label_num_to_disease_map.json') as f:\n    mapping = json.loads(f.read())\n    mapping = {int(k): v for k,v in mapping.items()}\n\nprint(mapping)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"<b>このdoctoeのコースでは、4つの病気について学ぶ必要がありますが、その前に、これらの病気の名前をデータセットのラベルにマッピングすることができます。 </b>"},{"metadata":{"trusted":true},"cell_type":"code","source":"train['label_names'] = train['label'].map(mapping)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Step 1.1 健康なもの(Healthy)を学ぶ\n\n今、私達は健康なイメージを見始め、健康なキャッサバの葉の特徴の私達の理解を形成することができる1つの場所ですべてを持っている。`次は Google からの健康なキャッサバの葉のイメージです` \n\n![画像](https://cdn.shortpixel.ai/client/to_avif,q_lossless,ret_img,w_795,h_532/https://organic.ng/wp-content/uploads/2017/02/CASSAVA-LEAF.jpg)\n\n* 上の画像から、健康的なキャッサバの葉の特徴の一つは、多くのカット、テクスチャの変更、黄色がかったグラデーション、等なしでかなり緑と直立する必要があることを言うことができます。\n\n次に、データセットの中の健康なものを見て、上の画像と密接に混ざっているかどうかを見てみましょう。"},{"metadata":{"trusted":true},"cell_type":"code","source":"train[train['label_names']=='Healthy']['image_id'].count()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* 21kの画像のうち、2577だけが健康的(Healthy)なものである、ラベルの不均衡は明らかに目に見えている"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_images(class_id=4, \n    label='Healthy',\n    images_number=6,verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_images(class_id=4, \n    label='Healthy',\n    images_number=6,verbose=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"上記の関数を3～4回実行して、毎回新しい画像が表示されるのを注意深く観察すると、次のようなことに気づくでしょう。\n* すべての画像が葉をクローズアップしているわけではありませんし、人間の目には葉がほとんど見えない木全体が写っている画像もありますし、葉よりも茎が多く写っている画像もあります。\n* さらに驚くべきことは、健康な葉の画像の中には感染しているように見えるものもあり、黄色や黄色がかったグラデーションのような色をしているものもあります。\n\n### 外れ値の調査 :  \nポイント2を調査するために、私は次のようなアイデアを持っています。\n\n* ここでのアイデアは、健全な画像をクラスタリングし、それぞれのクラスタを見て、外れ値クラスタと破損クラスタを見つけることができるかどうかを確認することです。\n* クラスタリングのための特徴量を生成するために Resnet18 を使用します。"},{"metadata":{"trusted":true},"cell_type":"code","source":"def extract_features(image_id, model):\n    file = BASE_DIR/'train_images'/image_id\n    # load the image as a 224x224 array\n    # load_imgはkerasの関数で出力はPILのインスタンス\n    img = load_img(file, target_size=(224,224))\n    # convert from 'PIL.Image.Image' to numpy array\n    img = np.array(img) \n    # reshape the data for the model reshape(num_of_samples, dim 1, dim 2, channels)\n    reshaped_img = img.reshape(1,224,224,3) \n    # prepare image for model\n    # Numpy配列を前処理（モデルに画像のデータを入れる際に必要な処理か）\n    imgx = preprocess_input(reshaped_img)\n    # get the feature vector\n    #引数で定めたモデルで予測する\n    features = model.predict(imgx, use_multiprocessing=True)\n    \n    return features","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = ResNet50()\nmodel = Model(inputs = model.inputs, outputs = model.layers[-2].output)\n\nhealthy = train[train['label']==4]\nhealthy['features'] = healthy['image_id'].progress_apply(lambda x:extract_features(x,model))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.summary()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#画像数*2048\nfeatures.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"features = np.array(healthy['features'].values.tolist()).reshape(-1,2048)\nimage_ids = np.array(healthy['image_id'].values.tolist())\n\n# Kmeansでクラスタリング(healtyの画像のみで)\nkmeans = KMeans(n_clusters=5,n_jobs=-1, random_state=22)\nkmeans.fit(features)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#クラスタリングした際のラベル\nkmeans.labels_","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"groups = {}\nfor file, cluster in zip(image_ids,kmeans.labels_):\n    if cluster not in groups.keys():\n        #groupsというdictに、keyをcluster,valueを画像のファイル名を集めたリストとして格納\n        groups[cluster] = []\n        groups[cluster].append(file)\n    else:\n        groups[cluster].append(file)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def view_cluster(cluster):\n    plt.figure(figsize = (25,25));\n    # gets the list of filenames for a cluster\n    files = groups[cluster]\n    # only allow up to 30 images to be shown at a time\n    if len(files) > 30:\n        print(f\"Clipping cluster size from {len(files)} to 25\")\n        start = np.random.randint(0,len(files))\n        files = files[start:start+25]\n    # plot each image in the cluster\n    for index, file in enumerate(files):\n        plt.subplot(5,5,index+1);\n        img = load_img(BASE_DIR/'train_images'/file)\n        img = np.array(img)\n        plt.imshow(img)\n        plt.title(file)\n        plt.axis('off')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"view_cluster(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* クラスター3のほとんどの外れ値をクラスター化することができ、それらを簡単に可視化することができました。\n\n* 葉が傷んでいたり、茶色い斑点があったり、健康ではないように見えるものもあります。\n\n同じトピックに対処する多数の議論があります。\n* https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198363 -- Wrong Labels\n* https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/199606 --  Quality of Labels\n\n訓練セットのノイズを心配する必要はありませんが、しかし、もしノイズがテストセットに含まれていて、ラベリングが同様に行われている場合はどうでしょうか？, それは問題かもしれません、我々は我々が確信するまで、訓練セットから何かを削除することはできません\n\n\nこのセクションを要約すると\n\n` 健康なキャッサバの葉の特徴`:\n* 主に緑色で、直立していて、茶色の斑点がほとんどない。\n* 黄色でも緑でも葉全体に均一な質感を与える"},{"metadata":{"trusted":true},"cell_type":"code","source":"view_cluster(2)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 病気を知る1：キャッサバ菌病（CBB）について\n\n今、私たちは健康なキャッサバの葉がどのように見えるかを知っているので、 最初の病気について学ぶために移動しましょう `CBBの症状`:\n\n* 黒い葉の斑点や病斑、角張った葉の斑点、若葉の萎凋による葉の早枯れや脱落、重度の攻撃などがあります。\n\n* 最初は葉脈によって制限された葉に角張った水浸しの斑点が発生し、葉の下の方にはっきりと見られます。斑点は急速に拡大し、特に葉縁に沿って合流し、褐色で黄色の縁取りをする（図1）。\n\n* 斑点の中心部にクリーム色の白色の液滴が発生し、その後、黄色に変化します。\n\n![図1](https://www.pestnet.org/fact_sheets/assets/image/cassava_bacterial_blight_173/thumbs/cassavabb_sml.jpg)\n![図2](https://www.pestnet.org/fact_sheets/assets/image/cassava_bacterial_blight_173/thumbs/cassavabb2_sml.jpg)\n\n\n詳細は [here](https://www.pestnet.org/fact_sheets/cassava_bacterial_blight_173.htm)"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_images(class_id=0, \n    label='CBB',\n    images_number=6,verbose=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* 症状の知識から、これらはCBB病に罹患していることは間違いありませんし、葉そのものではなく茎の画像を取得することは、病気の中には茎で判断できるものもありますので、茎の画像はノイズではないかもしれません。\n\n* IMG - '1926670152.jpg'のような画像のいくつかでは、茶色の斑点は非常に小さく、葉は健康なもののように見え、健康な画像の多くはまた、そのような小さな茶色を持っており、識別するのは難しいかもしれません\n\n* このカテゴリの病気の私の理解から、私はRandomCropping、コントラストの変化、任意の種類の色の変化は良いアイデアではないかもしれないと言うことができます。"},{"metadata":{},"cell_type":"markdown","source":"### 病気について学ぶ2：キャッサバグリーンモット（CGM)\n\n次の病気、 `Symptoms of CGM`に移ります\n\n* 葉に白斑が発生し、最初の小さな斑点から葉全体に広がり、葉緑素が失われていきます。若い葉は凹み、かすかな黄色の斑点が目立つ（図1）。(図 1)\n\n* この病気にかかると、葉に斑点状の症状が現れ、キャッサバモザイク病（CMD）の症状と混同されることがあります。重度のダメージを受けた葉は収縮し、乾燥して落ち、ローソク足のような特徴的な外観になります。(図2) (図 2)\n\n![](https://www.pestnet.org/fact_sheets/assets/image/cassava_green_mottle_068/thumbs/cgmv2_sml.jpg)\n![](https://www.pestnet.org/fact_sheets/assets/image/cassava_green_mottle_068/thumbs/cgmv_sml.jpg)\n\n詳細は [here](https://www.pestnet.org/fact_sheets/cassava_green_mottle_068.htm)"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_images(class_id=2, \n    label='CGM',\n    images_number=12,verbose=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### 推論\n\n* CGMの症状を読み、データセットの画像を見た後、CGMの葉、CBBの葉、健康な葉の違いを明確に伝えることができます。\n* CGMの葉はビエン(viens)に沿って葉にかすかに黄色の斑点があり、CBBの葉は茶色の斑点があり、健康な葉は完全に緑か完全に黄色である。\n* また、このクラスでもあまり外れた人はいません。\n\n\n### 病気を知る3：キャッサバモザイク病（CMD）について\n\n`CMDの症状`:\n\n* CMDは、モザイク、斑入り、葉の変形やねじれ、葉や植物のサイズの全体的な減少を含む様々な葉状の症状を生成します。\n\n\n* この病気によって影響を受けた葉は、通常の緑色のパッチを持ち、重症度に応じて黄色と白の異なる割合で混合されています。"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_images(class_id=3, \n    label='CMD',\n    images_number=6,verbose=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### 推論\n\n* 我々は、CGMとCMDは非常に近い症状を持っていることを見ることができ、また、かなり似たような画像を持っている、多くの場合、専門家は、これらのラベルを混乱させるかもしれません、我々はそれがモデルのためのものになりますどのように大きな課題を想像することができました。\n\n* このカテゴリーでも外れ者はないか、あるいは非常に少ないように思われます。"},{"metadata":{},"cell_type":"markdown","source":"### 病気を知る4：キャッサバ褐条病(CBSD)\n\n今、私は最後にこれを選んだ理由は、我々はこのカテゴリのために2つの異なる種類の画像を持っているためです。\n\n* 一つは、葉っぱ・植物の画像です。\n* もう一つは、結節性の根の画像ですが、これはジャガイモやノイズと誤解されやすいのですが、データセットにノイズが含まれていることから、この病気の識別に偏りが出てしまいます。したがって、データセットに写っている茶色くて不恰好なものは、カサベの結節性根であり、この病気もまた、これらの画像から識別することができることを明確にしておきます。\n\n今すぐ `CBSDの症状`を見てみましょう。\n\n* CBSDの葉の症状は、比較的大きな黄色のパッチを形成するために拡大し、合体するかもしれない特徴的な黄色または壊死性の静脈のバンディングで構成されています。\n* 塊根の症状は、塊茎内の黒褐色の壊死領域と根のサイズの減少で構成されています。\n\nこのカテゴリに存在する2つのタイプの画像と、それら2つの異なる画像に見られる症状を明確に理解することができましたので、データを見てみましょう。"},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_images(class_id=1, \n    label='CBSD',\n    images_number=12,verbose=1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* 筒状の根の画像のクラスターをデータから取り出せるか試してみよう"},{"metadata":{"trusted":true},"cell_type":"code","source":"CBSD = train[train['label']==1]\nCBSD['features'] = CBSD['image_id'].progress_apply(lambda x:extract_features(x,model))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"features_cbsd = np.array(CBSD['features'].values.tolist()).reshape(-1,2048)\nimage_ids_cbsd = np.array(CBSD['image_id'].values.tolist())\n\n# Clustering\nkmeans_cbsd = KMeans(n_clusters=5,n_jobs=-1, random_state=22)\nkmeans_cbsd.fit(features_cbsd)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"groups_cbsd = {}\nfor file, cluster in zip(image_ids_cbsd,kmeans_cbsd.labels_):\n    if cluster not in groups_cbsd.keys():\n        groups_cbsd[cluster] = []\n        groups_cbsd[cluster].append(file)\n    else:\n        groups_cbsd[cluster].append(file)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def view_cluster(cluster):\n    plt.figure(figsize = (25,25))\n    # gets the list of filenames for a cluster\n    files = groups_cbsd[cluster]\n    # only allow up to 30 images to be shown at a time\n    if len(files) > 30:\n        print(f\"Clipping cluster size from {len(files)} to 25\")\n        start = np.random.randint(0,len(files))\n        files = files[start:start+25]\n    # plot each image in the cluster\n    for index, file in enumerate(files):\n        plt.subplot(5,5,index+1);\n        img = load_img(BASE_DIR/'train_images'/file)\n        img = np.array(img)\n        plt.imshow(img)\n        plt.title(file)\n        plt.axis('off')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"view_cluster(4)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* 我々は正常に80画像の1つのクラスタに管状根画像をクラスタ化することができました、それゆえに今、我々はすべてのIDSを取得することができ、この情報を使用する方法についての様々なアイデアを考えるかもしれません"},{"metadata":{"trusted":true},"cell_type":"code","source":"view_cluster(3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 私たちの調査結果の概要：ステップ1の終了\n\n最初のEDAの結果をまとめてみましょう。\n\n* 健康な画像は正しくラベル付けされていない可能性があり、間違ってラベル付けされた画像は5つのクラスターのうち3つのクラスターにあります。\n* 完全に黄色の葉は、常に葉が潜在的な病気を持っていることを示していない可能性があります。\n* 葉に茶色の斑点があるのは、キャッサバのバクテリア・ベト病(CBB)を示しています。 \n* すべてのイメージに異なった背景およびスケールの変化があります \n* 画像は一日の異なる時間帯に撮影されているため、異なる照明と露出を持っています。\n* キャッサバグリーンモットル(CGM)とキャッサバモザイク病(CMD)は、画像と同様に非常に似た症状を持っており、簡単に互いに誤認表示される可能性があります。また、キャッサバモザイク病の例が13kもあるので、モデルがCGMをCGMとラベル付けする際に最もミスが多い可能性が高いです。\n* 一枚の画像/キャッサバの植物には複数の共起性疾患が含まれている可能性があります。モデルはラベル付けを混乱させる\n* CBSDはデータセットの中に2種類の画像を持っていますが、1つは植物/葉の画像で、もう1つはジャガイモやランダムノイズと誤解されやすい根の画像です。\n\n\nリーフドクターになるための最初のステップが完了した後、患者さんのことを理解し、様々な病気のことを理解することができました。このステップは、私たちはより良いソリューションを構築するためのユニークなソリューション/プランをデバイスに役立ちます。\n\n<b>注：私はより多くの発見を続けるように、私はこのセクションでより多くのそのような知見を追加していきます。</b>"},{"metadata":{},"cell_type":"markdown","source":"## データの重複  我々が見逃していたもの\n\n議論の場を調べていたら、データセットの中に画像が重複している可能性について話している[この](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198202) スレッドを見つけました。画像のデータセットの中に重複した画像があるという話はとても興味深いものです．\n\n* 画像の完全なコピーについて話している\n* 私たちは、特定の画像に似ている画像について話しています。例：画像1はトリミングされたか回転され、画像2として保存されています。\n\nさて、画像データセットの中から重複画像（正確なコピー）や類似画像を見つけて識別する方法がいくつかあります。私は画像ハッシュ化の方法を使い、 [ここ](https://www.kaggle.com/appian/let-s-find-out-duplicate-images-with-imagehash) で見つけたノートに従っていきます。"},{"metadata":{"trusted":true},"cell_type":"code","source":"funcs = [\n        imagehash.average_hash,\n        imagehash.phash,\n        imagehash.dhash,\n        imagehash.whash,\n    ]\n\nimage_ids = []\nhashes = []\n\nfor path in tqdm(glob.glob(str(BASE_DIR/'train_images'/'*.jpg' ))):\n    image = Image.open(path)\n    image_id = os.path.basename(path)\n    image_ids.append(image_id)\n    hashes.append(np.array([f(image).hash for f in funcs]).reshape(256))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"hashes_all = np.array(hashes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"hashes_all.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"numpy配列をトーチテンソルに変換して類似度計算を高速化します。"},{"metadata":{"trusted":true},"cell_type":"code","source":"hashes_all = torch.Tensor(hashes_all.astype(int)).cuda()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"すべての画像ペア間の類似度を計算します。値を256で割って正規化（0-1）します。"},{"metadata":{"trusted":true},"cell_type":"code","source":"%time sims = np.array([(hashes_all[i] == hashes_all).sum(dim=1).cpu().numpy()/256 for i in range(hashes_all.shape[0])])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sims.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"しきい値の設定"},{"metadata":{"trusted":true},"cell_type":"code","source":"indices1 = np.where(sims > 0.9)\nindices2 = np.where(indices1[0] != indices1[1])\nimage_ids1 = [image_ids[i] for i in indices1[0][indices2]]\nimage_ids2 = [image_ids[i] for i in indices1[1][indices2]]\ndups = {tuple(sorted([image_ids1,image_ids2])):True for image_ids1, image_ids2 in zip(image_ids1, image_ids2)}\nprint('found %d duplicates' % len(dups))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"重複した画像のプロット"},{"metadata":{"trusted":true},"cell_type":"code","source":"'''\ncode taken from https://www.kaggle.com/nakajima/duplicate-train-images?scriptVersionId=47295222\n'''\n\nduplicate_image_ids = sorted(list(dups))\n\nfig, axs = plt.subplots(2, 2, figsize=(15,15))\n\nfor row in range(2):\n        for col in range(2):\n            img_id = duplicate_image_ids[row][col]\n            img = Image.open(str(BASE_DIR/'train_images'/img_id))\n            label =str(train.loc[train['image_id'] == img_id].label.values[0])\n            axs[row, col].imshow(img)\n            axs[row, col].set_title(\"image_id : \"+ img_id + \"  label : \" + label)\n            axs[row, col].axis('off')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"重複を見つける方法は他にもあり、データセットの中にソフトな重複がある場合には、このカーネルの後のバージョンで提供されます。"},{"metadata":{},"cell_type":"markdown","source":"# ステップ2：方法論について学ぶ\n\nこんにちは、医師はあなたの2年目へようこそ、あなたの最終課題を完了するために、あなたは今、あなたの処分で持っているツールを理解し、それらを使用する方法を理解する必要があります、以下は、ツールを学ぶために従うべきステップバイステップのガイドです。\n\n* [Beginner Article](https://adeshpande3.github.io/adeshpande3.github.io/A-Beginner's-Guide-To-Understanding-Convolutional-Neural-Networks/)\n* [Course By Andrew NG](https://www.coursera.org/learn/convolutional-neural-networks)\n* [Applying CNNS using Keras and tensorflow](https://www.coursera.org/learn/convolutional-neural-networks-tensorflow)\n* [Course from Fast.ai](https://course.fast.ai/videos/?lesson=1)"},{"metadata":{},"cell_type":"markdown","source":"# ステップ3：最終プロジェクトの構築\n\n\nOhk now Docs , its time for you to build the final project . これは最終的なプロジェクトなので、誰もが自分自身で構築することを意味していますが、ここでは私が使用したものの要約を書き、さらにプロジェクトを改善するための方法を提案します。\n\n最後に、私はまた、競争の全体のコースの中で試すためのもの/外を見るためのものを追加します。\n\n`ベースラインモデルのまとめ`:\n\nこのモデルはキャッサバ2019大会の優勝解を元にしているので、できるだけ近い形で再現してみたいと思います。\n\n* SE-ResNext50\n* Dimension = (384,384)\n* Epochs = 10\n* Custom LR scheduler \n* Weights saved on best loss : Categorical CrossEntropy\n* Basic Augs : HorizontalFlip,VerticalFlip,Rotate,RandomBrightness,ShiftScaleRotate,cutout,centercrop,zoom,randomscale\n* No TTA（testデータにaugumentを行わない）\n\n<font color ='red' >GPUのためのkaggleに制限されているので、私の5つ折りモデルはまだ実行されているので、今のところはSeResNext50の事前学習された重みだけを使用しています。このノートブックは、異なる設定/アイデアで数回更新されますので、チューニングを維持してください</color>"},{"metadata":{},"cell_type":"markdown","source":"## 設定とユーティリティ機能"},{"metadata":{"trusted":true},"cell_type":"code","source":"DIM = (384,384)\n\nNUM_WORKERS = 12\nTEST_BATCH_SIZE = 16\nSEED = 2020\n\nDEVICE = \"cuda\"\n\nMEAN = [0.485, 0.456, 0.406]\nSTD = [0.229, 0.224, 0.225]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Augmentations"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_test_transforms():\n\n    return albumentations.Compose(\n        [albumentations.Normalize(MEAN, STD, max_pixel_value=255.0, always_apply=True),\n        ToTensorV2(p=1.0)\n        ]\n    )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Cassava Dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"class CassavaDataset(Dataset):\n    def __init__(self,image_ids,labels,dimension=None,augmentations=None):\n        super().__init__()\n        self.image_ids = image_ids\n        self.labels = labels\n        self.dim = dimension\n        self.augmentations = augmentations\n        \n    def __len__(self):\n        # len(上で定義したクラスのインスタンス)で返す値。\n        return len(self.image_ids)\n    \n    def __getitem__(self,idx):\n        # 上で定義したクラスのインスタンス[idx]で返す値\n        img = cv2.imread(str(BASE_DIR/'test_images'/self.image_ids[idx]))\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n                         \n        if self.dim:\n            img = cv2.resize(img,self.dim)\n        \n        if self.augmentations:\n            augmented = self.augmentations(image=img)\n            image = augmented['image']\n                         \n        return {\n            'image': image,\n            'target': torch.tensor(self.labels[idx],dtype=torch.float)\n        }","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Model : SE_Resnext50"},{"metadata":{"trusted":true},"cell_type":"code","source":"class CassavaModel(nn.Module):\n    def __init__(self, model_name='seresnext50_32x4d',out_features=5,pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_name, pretrained=pretrained)\n        \n        #resnet50のモデルに最後出力の次元を揃えるために1層追加している。\n        \n        n_features = self.model.last_linear.in_features\n        self.model.last_linear = nn.Linear(n_features, out_features)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Prediction Function Single Model"},{"metadata":{"trusted":true},"cell_type":"code","source":"def predict_single_model(data_loader,model,device):\n    model.eval()\n    tk0 = tqdm(enumerate(data_loader), total=len(data_loader))\n    fin_out = []\n    \n    with torch.no_grad():\n        \n        for bi, d in tk0:\n            images = d['image']\n            targets = d['target']\n            \n            images = images.to(device)\n            targets = targets.to(device)\n            \n            batch_size = images.shape[0]\n            \n            outputs = model(images)\n            \n            fin_out.append(F.softmax(outputs, dim=1).detach().cpu().numpy())\n            \n    return np.concatenate(fin_out)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Engine"},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_sub = pd.read_csv('../input/cassava-leaf-disease-classification/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def predict(weights):\n    '''\n    weights : List of paths in case of K fold model inference\n    '''\n    pred = np.zeros((len(sample_sub),5,5))\n    \n    # Defining DataSet\n    test_dataset = CassavaDataset(\n        image_ids=sample_sub['image_id'].values,\n        labels=sample_sub['label'].values,\n        augmentations=get_test_transforms(),\n        dimension = DIM\n    )\n    \n    test_loader = torch.utils.data.DataLoader(\n        test_dataset,\n        batch_size=TEST_BATCH_SIZE,\n        num_workers=NUM_WORKERS,\n        shuffle=False,\n        pin_memory=True,\n        drop_last=False,\n    )\n    \n    # Defining Device\n    device = torch.device(\"cpu\")\n    \n    for i,weight in enumerate(weights):\n        # Defining Model for specific fold\n        model = CassavaModel(out_features=5,pretrained=True)\n        \n        # loading weights\n        #model.load_state_dict(torch.load(weight))\n        model.to(device)\n        \n        #predicting\n        pred[:,:,i] = predict_single_model(test_loader,model,device)\n    \n    return pred","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Preparing Final Submission"},{"metadata":{"trusted":true},"cell_type":"code","source":"pred = predict([1])\nprint(pred)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pred = pred.mean(axis=-1)\nprint('Prediction Before Argmax',pred)\npred = pred.argmax(axis=1)\nprint('Final Prediction',pred)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_sub['label'] = pred\nsample_sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_sub.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# 結論\n\n競争が始まったばかりなので、試してみることがたくさんありますが、私はこのノートを更新してみます。\n\n私のノートを読んでくれてありがとう , 私はあなたがそれから有用な何かを得たことを願っています。"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}