{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":114397,"databundleVersionId":13696770,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Clustering the BioTrove Dataset: Baseline submission\n\nThis notebook generates a submission file without clustering anything. If the score of your clustering algorithm is lower, something has gone wrong...\n\nThis version of the notebook implements the following strategy:\n1. Family clusters: The metadata file gives us the family of every image; it's obvious that predicting exactly this family will give a perfect score.\n2. Genus clusters: The genera are only a little bit more granular than the families; we simply predict the families.\n3. Species clusters: We create as many clusters as there are images so that every image is alone in its own cluster.\n\nTo learn more about the competition metric (normalized mutual information score), study [this section of the scikit-learn user guide](https://scikit-learn.org/stable/modules/clustering.html#mutual-info-score).\n\nReference\n- [Kaggle competition](https://www.kaggle.com/competitions/biotrove-clustering)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-17T11:07:36.751584Z","iopub.execute_input":"2025-09-17T11:07:36.751880Z","iopub.status.idle":"2025-09-17T11:07:39.782450Z","shell.execute_reply.started":"2025-09-17T11:07:36.751856Z","shell.execute_reply":"2025-09-17T11:07:39.780995Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"metadata = pd.read_csv('/kaggle/input/biotrove-clustering/metadata.csv')\nsub = pd.DataFrame({\n    'hash_id': metadata.hash_id,\n    'family_cluster': metadata.family,\n    'genus_cluster': metadata.family,\n    'species_cluster': np.arange(len(metadata))\n})\ndisplay(sub)\nsub.to_csv('submission.csv', index=False)\nprint()\n!head submission.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-17T11:07:39.785003Z","iopub.execute_input":"2025-09-17T11:07:39.785513Z","iopub.status.idle":"2025-09-17T11:07:40.215983Z","shell.execute_reply.started":"2025-09-17T11:07:39.785482Z","shell.execute_reply":"2025-09-17T11:07:40.214889Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Lesson learned\n\n> **Before you start designing a model, understand the metric!**","metadata":{}}]}