{"cells":[{"metadata":{},"cell_type":"markdown","source":"1.plot_one_mask_GlobalAndLocal (SERIES, figsize = (14, 14), alpha = 0.35)  \nBasic visualization functions:  \nIn the whole image, a mask is visually highlighted, the background pixel value becomes smaller, and the contrast is increased.  \nTake a mask area from the original image for local visualization, and the background is 0 pixels.  \n\nUnderstand the data  \n2.plot_one_ClassID (dataframe, ClassId, nums = 10, figsize = (14, 14), alpha = 0.35)  \nGlobal and local visualization of different graphs of a certain ClassID. Take the first 10 pictures from the training set. The default arrangement is 10 rows and 3 columns. Each row is a different picture. The first column is the original image, the second column is the global mask, and the third column is the local area visualization. Some ClassIDs do not know what it is, you can understand it through this visualization.  \n\n3.plot_one_Attribute (dataframe, Attribute, figsize = (14,14), alpha = 0.35)  \nGlobal and local visualization of different ClassIDs of a certain Attribute. The default arrangement is n rows and 3 columns, each row is a picture of a different ClassId, the first column is the original image, the second column is the global mask, and the third column is the local area visualization. I want to understand the characteristics of attributes through this process.  \n    \n\nSome visualizations to understand properties such as length, vein, etc.    \nUnderstand the relationship between level2 and level1   \n\nTODU:  \nTest truth value comparison:  \n4.The visualization of two overlapping masks in one picture, used to compare the difference between prediction and truth mask.   "},{"metadata":{},"cell_type":"markdown","source":"1.plot_one_mask_GlobalAndLocal(SERIES, figsize=(14, 14) ,alpha = 0.35)  \n可视化基本功能：  \n在全图中 可视化突出一个掩模，背景像素值变小，增加对比度。  \n从原图中截取一个掩模的区域 进行局部可视化，背景为0像素。  \n\n了解数据   \n2.plot_one_ClassID(dataframe,ClassId,nums=10,figsize=(14, 14),alpha = 0.35 )   \n某一类ClassID 不同图的全局和局部可视化。从训练集中取前面10个图片，排列方式默认是10行3列，每一行是不同的图片，第一列是原图，第二列是全局掩模，第三列是局部区域可视化。有些ClassID不知道是什么东西，可以通过这个可视化了解下。     \n\n3.plot_one_Attribute(dataframe,Attribute,figsize = (14,14),alpha = 0.35)    \n某一类Attribute的不同ClassID的全局和局部可视化。排列方式默认是n行3列，每一行是不同的ClassId的图片，第一列是原图，第二列是全局掩模，第三列是局部区域可视化。想通过这个过程了解属性有怎样的特征。  \n    \n\n一些可视化搞明白长度、脉络等属性    \n搞明白level2和level1的关系  \n\nTODU:  \n测试真值对比：  \n4.一张图片中 两个有重叠掩模的可视化，用于比较预测和真值掩模的差异。    \n"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport matplotlib.image as mpimg\nimport numpy as np\nfrom matplotlib import pyplot as plt\nimport gc\nimport cv2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/imaterialist-fashion-2020-fgvc7/train.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def rle_to_mask(rle_string,height,width):\n    # https://www.kaggle.com/tanreinama/prediction-and-submission-of-attributes\n    rows, cols = height, width\n    if rle_string == -1:\n        return np.zeros((height, width))\n    else:\n        rleNumbers = [int(numstring) for numstring in rle_string.split(' ')]\n        rlePairs = np.array(rleNumbers).reshape(-1,2)\n        img = np.zeros(rows*cols,dtype=np.uint8)\n        for index,length in rlePairs:\n            index -= 1\n            img[index:index+length] = 255\n        img = img.reshape(cols,rows)\n        img = img.T\n        return img","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# plot_one_mask_GlobalAndLocal"},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_one_mask_GlobalAndLocal(SERIES, figsize=(14, 14) ,alpha = 0.35):\n    \n    mask = rle_to_mask(SERIES['EncodedPixels'],SERIES['Height'],SERIES['Width'])\n    image = cv2.imread(\"../input/imaterialist-fashion-2020-fgvc7/train/\"+str(SERIES['ImageId'])+\".jpg\")\n    b,g,r=cv2.split(image)\n    image = cv2.merge([r,g,b])\n    \n    assert image.shape[0:2] == mask.shape[0:2]\n    shape = image.shape[0:2]\n    fig, ax = plt.subplots(nrows=1, ncols=3, figsize=figsize)\n    ax[0].imshow(image)\n    ax[0].set_title('ImageId: '+SERIES['ImageId'])\n    \n    ax[1].imshow(image)\n    ax[1].imshow(mask, alpha=alpha) # 重叠 : overlapped\n    ax[0].axis('off')\n    ax[1].axis('off')\n    ax[1].set_title('ClassId: '+str(SERIES['ClassId']))\n\n    image[mask==0] = 255 # 背景为空白 : background is white\n    where = np.where(image < 255) # 取掩模最小区域 : minimum mask area \n    if len(where[0]) > 0 and len(where[1]) > 0:\n        y1,y2,x1,x2 = min(where[0]),max(where[0]),min(where[1]),max(where[1])\n    ax[2].imshow(image[y1:y2,x1:x2])\n    ax[2].set_title('AttributesIds: '+SERIES['AttributesIds'])\n    \n    ax[0].axis('off')\n    ax[1].axis('off')    \n    ax[2].axis('off')    \n    plt.show()\n    gc.collect()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# plot_one_ClassID(dataframe,ClassId,nums=10,figsize=(14, 14),alpha = 0.35 )"},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_one_ClassID(dataframe,ClassId,nums=10,figsize=(14, 14),alpha = 0.35 ):\n    # find\n    result = dataframe[dataframe.ClassId == ClassId][0:nums]    \n    \n    # plot\n    for index,ser in result.iterrows():\n        plot_one_mask_GlobalAndLocal(ser,figsize,alpha)    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_one_ClassID(train_df , 1)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# plot_one_Attribute(dataframe,Attribute,figsize = (14,14),alpha = 0.35)"},{"metadata":{"trusted":true},"cell_type":"code","source":"def plot_one_Attribute(dataframe,Attribute,figsize = (14,14),alpha = 0.35):\n    # find\n    # 筛选出有某个属性的样本 : Filter out samples with a certain attribute\n    # https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.str.contains.html#pandas.Series.str.contains\n    Attribute = str(Attribute)\n    AttributesId_sample = dataframe[dataframe['AttributesIds'].str.contains(Attribute, regex=False,na= False)] # na用来把NaN变为False : Na is used to make Nan false\n    \n    # 根据ClassId进行分组，每组取一个样本 : Group according to ClassId, one sample for each group.    \n    # https://pandas.pydata.org/pandas-docs/stable/reference/groupby.html\n    # as_index=False这样ClassId就不会被当成索引 : as_index = False so that ClassId will not be used as an index\n    result = AttributesId_sample.groupby(['ClassId'],as_index=False).first() \n\n    # plot\n    for index,ser in result.iterrows():\n        plot_one_mask_GlobalAndLocal(ser,figsize,alpha)    \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plot_one_Attribute(train_df , 317)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}