{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":71698,"databundleVersionId":7906362,"sourceType":"competition"}],"dockerImageVersionId":30673,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BugNist24 Training Data EDA and Characterization\nThis notebook is part of the BugNIST24 competition and being shared to accelerate new entrants understanding of the training data. There will be no modeling of the data here, only characterization, measurements, and visualizations of data. We will build up to a visualization with a computed center of  mass based on the projections. And then we will output the data to lists, collate those lists into a dataframe, and finally a csv for training use later.\n\nFeel Free to Fork this notebook, load as a precursor to your notebook, or use this code as you please. (To add this notebook as a source to your competition notebook click on \"Add Input\" in the right pane, search for \"BugNist24-Train-CSV-wCentersV1,\" and the plus sign next to the notebook.)\n\nIf you like the code, upvotes are appreciated.\n\n### What we'll do:\n1. Import Libraries\n2. Find all the files\n3. Show improper (flat tif) and then proper ways (tif volumes) to open the files.\n4. Build initial visualizer (incorrect solid)\n5. Build proper visualizer (marching cubes) with initial center point attempt.\n6. Visualize several bugs and center points to test hypothesis.\n7. Set up some lists to aid other's with dataloader or TF Dataset builds.\n8. Test and output computed data to CSV.\n\nLet's get started...","metadata":{}},{"cell_type":"code","source":"# import some libraries\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os # operating system control\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\nimport matplotlib.pyplot as plt # graphing\nfrom mpl_toolkits import mplot3d #matplot toolkit\nimport PIL as pil #Python Image Library\nimport re #regedit\nfrom skimage import io # in and out functions from a second image access library\nfrom skimage import measure # Used to define the marching cubes algorithm\nimport plotly.graph_objects as go #plotly image library\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-03-27T14:19:12.711071Z","iopub.execute_input":"2024-03-27T14:19:12.711579Z","iopub.status.idle":"2024-03-27T14:19:15.167675Z","shell.execute_reply.started":"2024-03-27T14:19:12.711544Z","shell.execute_reply":"2024-03-27T14:19:15.166057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # import open3d\n# !pip install open3d\n# import open3d as o3d \n# print(o3d.__version__)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:19:15.170472Z","iopub.execute_input":"2024-03-27T14:19:15.171170Z","iopub.status.idle":"2024-03-27T14:19:15.177486Z","shell.execute_reply.started":"2024-03-27T14:19:15.171125Z","shell.execute_reply":"2024-03-27T14:19:15.175934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's see if we can find the files to train...\nROOT = '/kaggle/input/bugnist2024fgvc/BugNIST_DATA/train/'\ni = 0\nfor dirname in os.listdir(ROOT):\n    print(dirname, len(os.listdir(ROOT + dirname + \"/\")))\n    i += len(os.listdir(ROOT + dirname + \"/\"))\nprint(\"Total Training Files:\", i)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:19:15.179496Z","iopub.execute_input":"2024-03-27T14:19:15.179941Z","iopub.status.idle":"2024-03-27T14:19:19.678544Z","shell.execute_reply.started":"2024-03-27T14:19:15.179896Z","shell.execute_reply":"2024-03-27T14:19:19.677067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cool, we found the 9154 files to train the models with. \n\nWhat does one of those look like? (we'll show how to do this wrong with PIL first)","metadata":{}},{"cell_type":"code","source":"for i in range(3):\n    im = pil.Image.open('/kaggle/input/bugnist2024fgvc/BugNIST_DATA/train/AC/bcrick_10_00' + str(i)  + '.tif')\n    plt.imshow(im)\n    plt.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:19:19.679996Z","iopub.execute_input":"2024-03-27T14:19:19.680382Z","iopub.status.idle":"2024-03-27T14:19:20.640307Z","shell.execute_reply.started":"2024-03-27T14:19:19.680352Z","shell.execute_reply":"2024-03-27T14:19:20.638911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# what does the raw image data look like?\nim = np.array(im)\nprint(\"Max:\", im.max() , \" Min:\", im.min(), \" Mean:\", im.mean())\nim[30,:]\n","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:19:20.643758Z","iopub.execute_input":"2024-03-27T14:19:20.644180Z","iopub.status.idle":"2024-03-27T14:19:20.654216Z","shell.execute_reply.started":"2024-03-27T14:19:20.644147Z","shell.execute_reply":"2024-03-27T14:19:20.652923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great, when we open these images as flat tif files. They look to be grayscale (only one channel) and the scale is probably 0-255 (an assumption because we see a max of 6 at this point). But this is all WRONG. \n\nThese files are 3D volumes (X, Y, Z) and this method only extracts 2 of the dimensions. We need to try another way...\n\nNext, we'll use skimage to open the files and see the difference...","metadata":{}},{"cell_type":"code","source":"# open the image\nimg = io.imread('/kaggle/input/bugnist2024fgvc/BugNIST_DATA/train/AC/bcrick_10_007.tif')\n\n# print the shape\nprint(\"Shape: \" , img.shape)\n\n# look at the raw data from a 5 pixel x 5 pixel x 5 pixel section in the middle of the image\nimg[35:40,35:40,35:40]","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:19:20.655498Z","iopub.execute_input":"2024-03-27T14:19:20.655894Z","iopub.status.idle":"2024-03-27T14:19:20.689108Z","shell.execute_reply.started":"2024-03-27T14:19:20.655865Z","shell.execute_reply":"2024-03-27T14:19:20.687936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's better. Now we have a 128 pixel deep, 64 pixel tall, and 64 pixel wide image to work with. \n\nBefore we move on, we'll see how to do that properly with the PIL library for completeness.","metadata":{}},{"cell_type":"code","source":"dataset = pil.Image.open('/kaggle/input/bugnist2024fgvc/BugNIST_DATA/train/AC/bcrick_10_007.tif')\n\n# set height and width\nh,w = np.shape(dataset)\n\n# set number of pixels in depth\ntiffarray = np.zeros((dataset.n_frames,h,w))\n\n#loop through depth and put the data in place of the zeros\nfor i in range(dataset.n_frames):\n    dataset.seek(i)\n    tiffarray[i,:,:] = np.array(dataset)\n\nimg = tiffarray.astype(np.double);\nprint(img.shape)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:19:20.690945Z","iopub.execute_input":"2024-03-27T14:19:20.691392Z","iopub.status.idle":"2024-03-27T14:19:20.905868Z","shell.execute_reply.started":"2024-03-27T14:19:20.691349Z","shell.execute_reply":"2024-03-27T14:19:20.904942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have shown how to get the data out of the image two different ways and populate the arrays needed for later steps. The next order of business is to visualize the data. What do these things look like?","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize = (7, 7))\nax = fig.add_subplot(111, projection = '3d')\n\nax.voxels(img, alpha = 0.4)\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:19:20.907366Z","iopub.execute_input":"2024-03-27T14:19:20.907695Z","iopub.status.idle":"2024-03-27T14:20:20.772581Z","shell.execute_reply.started":"2024-03-27T14:19:20.907668Z","shell.execute_reply":"2024-03-27T14:20:20.771651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is good, we can see the image in 3D. \n\nThis is bad, that took forever and that isn't a bug. We are seeing the whole of the data clumped together. In order to visualize the data inside we'll need to try something different.\n\nNext we'll take a look at the marching cubes algorithm and visualize the results in plotly. We set the threshold (level) fairly high so that we eliminate clutter.","metadata":{}},{"cell_type":"code","source":"def plot_surface(img,draw=True):\n    \"\"\"\n    This function calculates the surfaces (using marching cubes), and displays them in 3D\n    \"\"\"\n    # use marching cubes to extract the surface, play around with the level value to see which level yields the best results\n    verts, faces, _, _ = measure.marching_cubes(img, level = 42)\n\n    # compute the faces and vertices \n    dats = go.Mesh3d(x = verts[:, 0], y = verts[:, 1], z = verts[:, 2],\n                                     i = faces[:, 0], j = faces[:, 1], k = faces[:, 2],\n                                     opacity = 0.5)\n    \n    #extract the points and the means\n    x_data = dats.x\n    y_data = dats.y\n    z_data = dats.z\n    \n    cenx, ceny, cenz = np.mean(x_data),np.mean(y_data),np.mean(z_data)\n    center_point = (cenx,ceny,cenz)\n    if draw:\n        # plot\n        fig = go.Figure(data = dats)\n        fig.update_layout(scene = dict(xaxis = dict(title='Z'),\n                                     yaxis = dict(title='Y'),\n                                     zaxis = dict(title='X')),\n                          margin = dict(l = 0, r  =0, t = 0, b = 0),\n                          title = \"Bug Projection\")\n\n        # Add a scatter plot for the center point\n        new_point = go.Scatter3d(\n        x= [cenx],  # X coordinate of the new point\n        y= [ceny],  # Y coordinate of the new point\n        z= [cenz],  # Z coordinate of the new point\n        mode='markers',\n        marker=dict(\n            size=10,\n            color='red',\n        ),\n        name='Center Mass'\n        )\n\n        # Add the new point to the figure\n        fig.add_trace(new_point)\n\n        fig.show()\n        print(\"Center of Mass Computed at: (Z,Y,X)\",center_point)\n    return [cenx,ceny,cenz] #remember x,y,z is plotly center --> Z,Y,X is our numpy output\nz,y,x = plot_surface(img)","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:20:20.773886Z","iopub.execute_input":"2024-03-27T14:20:20.774819Z","iopub.status.idle":"2024-03-27T14:20:21.549687Z","shell.execute_reply.started":"2024-03-27T14:20:20.774779Z","shell.execute_reply":"2024-03-27T14:20:21.548320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"skimage centroid computed at: (Z,Y,X)\", measure.centroid(img))\n# NOTICE THE DIFFERENCE!","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:20:21.551390Z","iopub.execute_input":"2024-03-27T14:20:21.551791Z","iopub.status.idle":"2024-03-27T14:20:21.578945Z","shell.execute_reply.started":"2024-03-27T14:20:21.551759Z","shell.execute_reply":"2024-03-27T14:20:21.577749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's better. Now we can see an image of a bug. We can discern what kind of bug it is. And we have the ability to pan, zoom, rotate, and measure. This projection will be very helpful when we start comparing the center mass of the bug to the predictions from our models. We have also computed the center mass in the process. We'll need that center mass for our models and training data...\n\n**Something to be careful of:** The centroid function in skimage.measure gives us the center point of the mass of the whole image. We want the center point of the projection. We'll have to compute the center mass off of each projection as part of our training.\n\nNow let's look at the 12th file (sorted) in each of the folders to see some bugs...","metadata":{}},{"cell_type":"code","source":"ROOT = '/kaggle/input/bugnist2024fgvc/BugNIST_DATA/train/'\nfolders = os.listdir(ROOT)\nfor i in folders:\n    files = sorted(os.listdir(ROOT+i))\n    img = io.imread(ROOT+i+'/'+files[12])\n    print('-----------------------------------------------------------')\n    print(ROOT+i+'/'+files[12])\n    z,y,x = plot_surface(img)\n","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:20:21.580710Z","iopub.execute_input":"2024-03-27T14:20:21.581918Z","iopub.status.idle":"2024-03-27T14:20:22.541759Z","shell.execute_reply.started":"2024-03-27T14:20:21.581872Z","shell.execute_reply":"2024-03-27T14:20:22.540517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You may have to rotate and spin the files to see, but the red dots are generally in the center of each plotted mass. The plotting is fairly fast (12 plots faster than the plotted cylinder above), so this will work to build our training data. We have the files ordered (sorted) properly because we see that each filename corresponds. And we are ready to move on to making some lists...\n\nWe'll use a pandas dataframe to hold all this data, and output to a csv. We'll keep the filename, the folder reference (discerns the type of bug), the center coordinates (z,y,x), and process all the images.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# empty lists\nfi, fo, fz, fy, fx, fif = [] , [], [], [] , [] , [] \n# progress counter\nk = 1\n\n# set location\nROOT = '/kaggle/input/bugnist2024fgvc/BugNIST_DATA/train/'\nfolders = os.listdir(ROOT)\n\n# loop through the folders\nfor i in folders:\n    files = sorted(os.listdir(ROOT+i))\n    #loop through the files\n    for j in files:\n        # addend file info\n        fi.append(ROOT+i+'/'+j)\n        fif.append(j)\n        fo.append(i)\n        \n        # get the file\n        img = io.imread(ROOT+i+'/'+j)\n        #compute the center\n        z,y,x = plot_surface(img,False)\n        #addend the z,y,x point to lists\n        fz.append(z)\n        fy.append(y)\n        fx.append(x)\n        # increment counter\n        k +=1\n        # report if divisible by 1000\n        if k % 1000 == 0:\n            print(k, \" Completed\")\n            \n        \n","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:20:22.543743Z","iopub.execute_input":"2024-03-27T14:20:22.544523Z","iopub.status.idle":"2024-03-27T14:26:36.009158Z","shell.execute_reply.started":"2024-03-27T14:20:22.544485Z","shell.execute_reply":"2024-03-27T14:26:36.007998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output = pd.DataFrame({\"FileLoc\":fi,\n                       \"Fildname\":fif,\n                       \"BugType\":fo,\n                       \"CenterZ\":fz,\n                       \"CenterY\":fy,\n                       \"CenterX\":fx})\n\n#test it\nprint(output.iloc[1,0])\noutput.head()","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:26:36.010699Z","iopub.execute_input":"2024-03-27T14:26:36.011136Z","iopub.status.idle":"2024-03-27T14:26:36.061183Z","shell.execute_reply.started":"2024-03-27T14:26:36.011102Z","shell.execute_reply":"2024-03-27T14:26:36.059929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test it again\nprint(output.iloc[9153,0])\noutput.tail()","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:26:36.065026Z","iopub.execute_input":"2024-03-27T14:26:36.065367Z","iopub.status.idle":"2024-03-27T14:26:36.081978Z","shell.execute_reply.started":"2024-03-27T14:26:36.065340Z","shell.execute_reply":"2024-03-27T14:26:36.080656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#test it again\nprint(output.iloc[3712,0])\noutput.iloc[3712:3722,:]","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:26:36.083354Z","iopub.execute_input":"2024-03-27T14:26:36.083713Z","iopub.status.idle":"2024-03-27T14:26:36.106421Z","shell.execute_reply.started":"2024-03-27T14:26:36.083684Z","shell.execute_reply":"2024-03-27T14:26:36.105083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.to_csv('/kaggle/working/BugNISTcentersV1.csv')","metadata":{"execution":{"iopub.status.busy":"2024-03-27T14:26:36.107719Z","iopub.execute_input":"2024-03-27T14:26:36.108202Z","iopub.status.idle":"2024-03-27T14:26:36.218321Z","shell.execute_reply.started":"2024-03-27T14:26:36.108164Z","shell.execute_reply":"2024-03-27T14:26:36.217095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And there we have it. A CSV of the data is output to the /kaggle/working folder. The CSV is ready for you to import into your notebook and continue our progress. Feel free to use this notebook as a data source in your notebook, I'll leave it public... (Under add data in the right pane, search for 'BugNist24-Train-CSV-wCentersV1' and click add.)\n\nIf you have ideas for improvements, I'd love to learn. \nI'll leave the comments open.\n\nHappy Kaggling!","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}