{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Petals to the Metal — baseline (transfer learning)\n\nM10-T002 baseline for the\n[Petals to the Metal — Flower Classification on TPU](https://www.kaggle.com/competitions/tpu-getting-started)\ncompetition.\n\n- **Backbone:** `MobileNetV2` pretrained on ImageNet (small, fast — sanity baseline, not a leaderboard chaser).\n- **Image size:** 192x192 (smallest published variant).\n- **Optimizer / epochs:** Adam (default LR), 3 epochs for local sanity (TPU kernel may train longer; see `EPOCHS` constant).\n- **Seed:** 42.\n- **Split:** Kaggle ships `train/` and `val/` TFRecord splits separately — we use them as-is (holdout accuracy on `val/`).\n\nWhen executed on Kaggle's kernel runner, this notebook reads TFRecords from the GCS path exposed by\n`kaggle_datasets.KaggleDatasets().get_gcs_path('tpu-getting-started')`. When executed locally\n(e.g. via `jupyter nbconvert --to notebook --execute`), it falls back to the local slice at\n`data/tpu-getting-started/local-slice/{train,val}/*.tfrec` that M10-T002 staged for sanity.\n\n**Defense-in-depth:** this notebook does NOT call `kaggle kernels push` from any cell. Kernel pushing\nhappens via `kca kaggle push` in M10-T004, never from inside the notebook itself.","metadata":{}},{"cell_type":"code","source":"import collections\nimport random\nimport time\nfrom pathlib import Path\n\nimport numpy as np\nimport tensorflow as tf","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:08.004886Z","iopub.execute_input":"2026-09-27T11:11:08.005128Z","iopub.status.idle":"2026-09-27T11:11:25.243678Z","shell.execute_reply.started":"2026-09-27T11:11:08.005096Z","shell.execute_reply":"2026-09-27T11:11:25.24304Z"}},"outputs":[{"name":"stderr","text":"2026-09-27 11:11:09.992355: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:467] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered\nWARNING: All log messages before absl::InitializeLog() is called are written to STDERR\nE0000 00:00:1790507470.254141      58 cuda_dnn.cc:8579] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered\nE0000 00:00:1790507470.325212      58 cuda_blas.cc:1407] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered\nW0000 00:00:1790507470.908669      58 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.\nW0000 00:00:1790507470.908707      58 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.\nW0000 00:00:1790507470.908710      58 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.\nW0000 00:00:1790507470.908712      58 computation_placer.cc:177] computation placer already registered. Please check linkage and avoid linking the same target more than once.\n","output_type":"stream"}],"execution_count":1},{"cell_type":"code","source":"print (\"import collections\")\nprint (\"import random\")\nprint (\"import time\")\nprint (\"from pathlib import Path\")\n\nprint (\"import numpy as np\")\nprint (\"import tensorflow as tf\")\n\n# Check your answer\n# q1.check()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.246237Z","iopub.execute_input":"2026-09-27T11:11:25.246747Z","iopub.status.idle":"2026-09-27T11:11:25.251771Z","shell.execute_reply.started":"2026-09-27T11:11:25.24671Z","shell.execute_reply":"2026-09-27T11:11:25.250985Z"}},"outputs":[{"name":"stdout","text":"import collections\nimport random\nimport time\nfrom pathlib import Path\nimport numpy as np\nimport tensorflow as tf\n","output_type":"stream"}],"execution_count":2},{"cell_type":"markdown","source":"## Configuration","metadata":{}},{"cell_type":"code","source":"SEED = 42\nIMAGE_SIZE = (192, 192)\nNUM_CLASSES = 104\nBATCH_SIZE = 32  # small for local CPU; Kaggle kernel can override\nEPOCHS = 3  # local sanity: 3; Kaggle kernel runner may bump (see KERNEL_EPOCHS)\nKERNEL_EPOCHS = 5  # used when running on Kaggle infra (see strategy block below)\nCOMPETITION = \"tpu-getting-started\"\n\nrandom.seed(SEED)\nnp.random.seed(SEED)\ntf.random.set_seed(SEED)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.25287Z","iopub.execute_input":"2026-09-27T11:11:25.253147Z","iopub.status.idle":"2026-09-27T11:11:25.277796Z","shell.execute_reply.started":"2026-09-27T11:11:25.253115Z","shell.execute_reply":"2026-09-27T11:11:25.277025Z"}},"outputs":[],"execution_count":3},{"cell_type":"code","source":"print (\"SEED\")\nprint (\"IMAGE_SIZE\")\nprint (\"NUM_CLASSES\")\nprint (\"BATCH_SIZE\")\nprint (\"EPOCHS\")\nprint (\"KERNEL_EPOCHS\")\nprint (\"COMPETITION\")\n\nprint (\"random.seed\")\nprint (\"np.random.seed\")\nprint (\"tf.random.set_seed\")\nprint (\"np.random.seed\")\n\n# Check your answer\n# q1.check()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.278758Z","iopub.execute_input":"2026-09-27T11:11:25.279316Z","iopub.status.idle":"2026-09-27T11:11:25.295978Z","shell.execute_reply.started":"2026-09-27T11:11:25.27929Z","shell.execute_reply":"2026-09-27T11:11:25.295059Z"}},"outputs":[{"name":"stdout","text":"SEED\nIMAGE_SIZE\nNUM_CLASSES\nBATCH_SIZE\nEPOCHS\nKERNEL_EPOCHS\nCOMPETITION\nrandom.seed\nnp.random.seed\ntf.random.set_seed\nnp.random.seed\n","output_type":"stream"}],"execution_count":4},{"cell_type":"markdown","source":"## Resolve data paths\n\nTry Kaggle's GCS-backed dataset path first (kernel runner). Fall back to the local M10-T002 slice.","metadata":{}},{"cell_type":"code","source":"def _resolve_gcs_path() -> str | None:\n    try:\n        from kaggle_datasets import KaggleDatasets  # type: ignore\n\n        return KaggleDatasets().get_gcs_path(COMPETITION)\n    except Exception as exc:\n        print(f\"[data-resolve] kaggle_datasets GCS path unavailable: {exc!r}\")\n        return None\n\n\ndef _resolve_kaggle_input_path() -> str | None:\n    # Try the direct competition mount first, then the /competitions/<slug>/ layout\n    # Kaggle uses for \"official\" Getting Started competitions (v3 diagnostic showed\n    # /kaggle/input had only ['competitions'], not the slug directly).\n    candidates = [\n        Path(\"/kaggle/input\") / COMPETITION,\n        Path(\"/kaggle/input/competitions\") / COMPETITION,\n    ]\n    for p in candidates:\n        if p.exists():\n            print(f\"[data-resolve] /kaggle/input mount found at {p}\")\n            return str(p)\n    # Diagnostic: list what IS under /kaggle/input/ so a future failure shows\n    # the real layout instead of an opaque \"not found\".\n    root = Path(\"/kaggle/input\")\n    if root.exists():\n        try:\n            print(\n                f\"[data-resolve] /kaggle/input contents: {sorted(p.name for p in root.iterdir())}\"\n            )\n            comp_root = root / \"competitions\"\n            if comp_root.exists():\n                print(\n                    f\"[data-resolve] /kaggle/input/competitions contents: \"\n                    f\"{sorted(p.name for p in comp_root.iterdir())}\"\n                )\n        except Exception as exc:\n            print(f\"[data-resolve] /kaggle/input listdir failed: {exc!r}\")\n    else:\n        print(\"[data-resolve] /kaggle/input does not exist\")\n    return None\n\n\ndef _resolve_local_slice() -> str | None:\n    # When this notebook is executed locally from the repo root.\n    for candidate in [\n        Path.cwd() / \"data\" / COMPETITION / \"local-slice\",\n        Path(\"/home/harry/test/tt_pangu/.claude/worktrees/m10-petals-to-the-metal/data\")\n        / COMPETITION\n        / \"local-slice\",\n    ]:\n        if candidate.exists():\n            return str(candidate)\n    return None\n\n\ndef _has_tfrecord_subdir(root: str, subdir: str) -> bool:\n    # Validate that a resolver-returned `root` actually contains the expected\n    # TFRecord layout. v4 errored because GCS resolver returned a path that\n    # didn't contain the shards (Kaggle TPU shortcut to the wrong mount).\n    try:\n        check = f\"{root}/{subdir}/train\"\n        return tf.io.gfile.exists(check)\n    except Exception:\n        return False\n\n\n_SUBDIR = f\"tfrecords-jpeg-{IMAGE_SIZE[0]}x{IMAGE_SIZE[1]}\"\n\n# Order matters: prefer the explicit /kaggle/input/competitions/<slug>/ mount\n# (validated by _has_tfrecord_subdir) over the GCS shortcut, then fall back to\n# the GCS path, then a local slice for off-Kaggle execution.\nDATA_ROOT, DATA_SUBDIR, RUNTIME = None, None, None\nfor label, candidate in (\n    (\"kaggle-input\", _resolve_kaggle_input_path()),\n    (\"kaggle-tpu-gcs\", _resolve_gcs_path()),\n):\n    if candidate and _has_tfrecord_subdir(candidate, _SUBDIR):\n        DATA_ROOT, DATA_SUBDIR, RUNTIME = candidate, _SUBDIR, label\n        print(f\"[data-resolve] accepted {label} root={candidate}\")\n        break\n    elif candidate:\n        print(\n            f\"[data-resolve] rejected {label} root={candidate} — \"\n            f\"{_SUBDIR}/train not present\"\n        )\n\nif DATA_ROOT is None:\n    local = _resolve_local_slice()\n    if local:\n        DATA_ROOT, DATA_SUBDIR, RUNTIME = local, \"\", \"local-slice\"\n\nif DATA_ROOT is None:\n    raise RuntimeError(\n        \"No data source resolved: no Kaggle mount contains \"\n        f\"{_SUBDIR}/train, and no local slice is present. \"\n        \"Run scripts/m10_t002_download_slice.py to stage a local slice.\"\n    )\n\nprint(f\"runtime={RUNTIME} data_root={DATA_ROOT} subdir={DATA_SUBDIR!r}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.297204Z","iopub.execute_input":"2026-09-27T11:11:25.297516Z","iopub.status.idle":"2026-09-27T11:11:25.511911Z","shell.execute_reply.started":"2026-09-27T11:11:25.297483Z","shell.execute_reply":"2026-09-27T11:11:25.511074Z"}},"outputs":[{"name":"stdout","text":"[data-resolve] /kaggle/input mount found at /kaggle/input/competitions/tpu-getting-started\n[data-resolve] kaggle_datasets GCS path unavailable: BackendError('Unexpected response from the service. Response: {\\'errors\\': [\"Dataset not found for directory \\'tpu-getting-started\\', please make sure you are passing a valid directory name under /kaggle/input\"], \\'error\\': {\\'code\\': 5}, \\'wasSuccessful\\': False}.')\n[data-resolve] accepted kaggle-input root=/kaggle/input/competitions/tpu-getting-started\nruntime=kaggle-input data_root=/kaggle/input/competitions/tpu-getting-started subdir='tfrecords-jpeg-192x192'\n","output_type":"stream"}],"execution_count":5},{"cell_type":"markdown","source":"## TFRecord decoding","metadata":{}},{"cell_type":"code","source":"LABELED_TFREC_FORMAT = {\n    \"image\": tf.io.FixedLenFeature([], tf.string),\n    \"class\": tf.io.FixedLenFeature([], tf.int64),\n}\n\nUNLABELED_TFREC_FORMAT = {\n    \"image\": tf.io.FixedLenFeature([], tf.string),\n    \"id\": tf.io.FixedLenFeature([], tf.string),\n}\n\n\ndef decode_image(image_bytes):\n    image = tf.image.decode_jpeg(image_bytes, channels=3)\n    image = tf.image.resize(image, IMAGE_SIZE)\n    image = tf.cast(image, tf.float32) / 255.0\n    image = tf.reshape(image, [*IMAGE_SIZE, 3])\n    return image\n\n\ndef read_labeled_tfrecord(example):\n    parsed = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    return decode_image(parsed[\"image\"]), tf.cast(parsed[\"class\"], tf.int32)\n\n\ndef read_unlabeled_tfrecord(example):\n    parsed = tf.io.parse_single_example(example, UNLABELED_TFREC_FORMAT)\n    return decode_image(parsed[\"image\"]), parsed[\"id\"]\n\n\ndef _split_glob(split: str) -> str:\n    if DATA_SUBDIR:\n        return f\"{DATA_ROOT}/{DATA_SUBDIR}/{split}/*.tfrec\"\n    return f\"{DATA_ROOT}/{split}/*.tfrec\"\n\n\ndef load_dataset(split: str, labeled: bool):\n    files = tf.io.gfile.glob(_split_glob(split))\n    if not files:\n        raise RuntimeError(\n            f\"No TFRecords found for split={split} glob={_split_glob(split)}\"\n        )\n    ds = tf.data.TFRecordDataset(files, num_parallel_reads=tf.data.AUTOTUNE)\n    ds = ds.with_options(tf.data.Options())\n    ds = ds.map(\n        read_labeled_tfrecord if labeled else read_unlabeled_tfrecord,\n        num_parallel_calls=tf.data.AUTOTUNE,\n    )\n    return ds\n\n\ndef get_training_dataset():\n    ds = load_dataset(\"train\", labeled=True)\n    ds = ds.repeat()\n    ds = ds.shuffle(2048, seed=SEED)\n    ds = ds.batch(BATCH_SIZE, drop_remainder=True)\n    ds = ds.prefetch(tf.data.AUTOTUNE)\n    return ds\n\n\ndef get_validation_dataset():\n    ds = load_dataset(\"val\", labeled=True)\n    ds = ds.batch(BATCH_SIZE)\n    ds = ds.prefetch(tf.data.AUTOTUNE)\n    return ds\n\n\ndef get_test_dataset():\n    ds = load_dataset(\"test\", labeled=False)\n    ds = ds.batch(BATCH_SIZE)\n    ds = ds.prefetch(tf.data.AUTOTUNE)\n    return ds","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.512806Z","iopub.execute_input":"2026-09-27T11:11:25.512992Z","iopub.status.idle":"2026-09-27T11:11:25.523162Z","shell.execute_reply.started":"2026-09-27T11:11:25.512974Z","shell.execute_reply":"2026-09-27T11:11:25.522276Z"}},"outputs":[],"execution_count":6},{"cell_type":"markdown","source":"## Dataset sizing\n\nPetals to the Metal's TFRecord filenames embed the per-shard count (e.g. `00-192x192-798.tfrec` = 798 examples).\nWe extract that to know `steps_per_epoch` without iterating the whole dataset.","metadata":{}},{"cell_type":"code","source":"def count_data_items(filenames):\n    import re\n\n    total = 0\n    for fn in filenames:\n        m = re.search(r\"-(\\d+)\\.tfrec$\", fn)\n        if m:\n            total += int(m.group(1))\n    return total\n\n\ntrain_files = tf.io.gfile.glob(_split_glob(\"train\"))\nval_files = tf.io.gfile.glob(_split_glob(\"val\"))\n\nNUM_TRAINING_IMAGES = count_data_items(train_files)\nNUM_VALIDATION_IMAGES = count_data_items(val_files)\nSTEPS_PER_EPOCH = max(1, NUM_TRAINING_IMAGES // BATCH_SIZE)\n\nprint(\n    f\"train shards={len(train_files)} train_images={NUM_TRAINING_IMAGES} \"\n    f\"val shards={len(val_files)} val_images={NUM_VALIDATION_IMAGES} \"\n    f\"steps_per_epoch={STEPS_PER_EPOCH}\"\n)\n\n# Fail loudly if the resolver pointed at an empty path — v3 silently trained on\n# zero shards because we trusted the resolver's word. Catch that here.\nif len(train_files) == 0 or len(val_files) == 0:\n    raise RuntimeError(\n        f\"Zero TFRecords resolved at {DATA_ROOT}/{DATA_SUBDIR or '<root>'}/{{train,val}}/*.tfrec — \"\n        f\"train_files={len(train_files)} val_files={len(val_files)}. \"\n        f\"The resolver returned a path that does not contain the expected shards; \"\n        f\"investigate the /kaggle/input layout (see earlier [data-resolve] prints) \"\n        f\"before pushing again.\"\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.525332Z","iopub.execute_input":"2026-09-27T11:11:25.526038Z","iopub.status.idle":"2026-09-27T11:11:25.669239Z","shell.execute_reply.started":"2026-09-27T11:11:25.526014Z","shell.execute_reply":"2026-09-27T11:11:25.668451Z"}},"outputs":[{"name":"stdout","text":"train shards=16 train_images=12753 val shards=16 val_images=3712 steps_per_epoch=398\n","output_type":"stream"}],"execution_count":7},{"cell_type":"markdown","source":"## Strategy (TPU vs GPU/CPU)","metadata":{}},{"cell_type":"code","source":"try:\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect()\n    strategy = tf.distribute.TPUStrategy(tpu)\n    print(f\"strategy=TPU replicas={strategy.num_replicas_in_sync}\")\n    EPOCHS = KERNEL_EPOCHS\nexcept (ValueError, tf.errors.NotFoundError, Exception):\n    strategy = tf.distribute.get_strategy()\n    print(f\"strategy=default replicas={strategy.num_replicas_in_sync}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.670146Z","iopub.execute_input":"2026-09-27T11:11:25.670447Z","iopub.status.idle":"2026-09-27T11:11:25.676149Z","shell.execute_reply.started":"2026-09-27T11:11:25.670413Z","shell.execute_reply":"2026-09-27T11:11:25.675282Z"}},"outputs":[{"name":"stdout","text":"strategy=default replicas=1\n","output_type":"stream"}],"execution_count":8},{"cell_type":"markdown","source":"## Model","metadata":{}},{"cell_type":"code","source":"def build_model():\n    base = tf.keras.applications.MobileNetV2(\n        input_shape=[*IMAGE_SIZE, 3],\n        include_top=False,\n        weights=\"imagenet\",\n    )\n    base.trainable = False\n    model = tf.keras.Sequential(\n        [\n            base,\n            tf.keras.layers.GlobalAveragePooling2D(),\n            tf.keras.layers.Dense(NUM_CLASSES, activation=\"softmax\"),\n        ],\n        name=\"petals_baseline_mobilenetv2\",\n    )\n    model.compile(\n        optimizer=tf.keras.optimizers.Adam(),\n        loss=tf.keras.losses.SparseCategoricalCrossentropy(),\n        metrics=[tf.keras.metrics.SparseCategoricalAccuracy(name=\"accuracy\")],\n    )\n    return model\n\n\nwith strategy.scope():\n    model = build_model()\n\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:25.677008Z","iopub.execute_input":"2026-09-27T11:11:25.677281Z","iopub.status.idle":"2026-09-27T11:11:31.10316Z","shell.execute_reply.started":"2026-09-27T11:11:25.677247Z","shell.execute_reply":"2026-09-27T11:11:31.102543Z"}},"outputs":[{"name":"stderr","text":"I0000 00:00:1790507487.551263      58 gpu_device.cc:2019] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 13756 MB memory:  -> device: 0, name: Tesla T4, pci bus id: 0000:00:04.0, compute capability: 7.5\nI0000 00:00:1790507487.557873      58 gpu_device.cc:2019] Created device /job:localhost/replica:0/task:0/device:GPU:1 with 13756 MB memory:  -> device: 1, name: Tesla T4, pci bus id: 0000:00:05.0, compute capability: 7.5\n","output_type":"stream"},{"name":"stdout","text":"Downloading data from https://storage.googleapis.com/tensorflow/keras-applications/mobilenet_v2/mobilenet_v2_weights_tf_dim_ordering_tf_kernels_1.0_192_no_top.h5\n\u001b[1m9406464/9406464\u001b[0m \u001b[32m━━━━━━━━━━━━━━━━━━━━\u001b[0m\u001b[37m\u001b[0m \u001b[1m1s\u001b[0m 0us/step\n","output_type":"stream"},{"output_type":"display_data","data":{"text/plain":"\u001b[1mModel: \"petals_baseline_mobilenetv2\"\u001b[0m\n","text/html":"<pre style=\"white-space:pre;overflow-x:auto;line-height:normal;font-family:Menlo,'DejaVu Sans Mono',consolas,'Courier New',monospace\"><span style=\"font-weight: bold\">Model: \"petals_baseline_mobilenetv2\"</span>\n</pre>\n"},"metadata":{}},{"output_type":"display_data","data":{"text/plain":"┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓\n┃\u001b[1m \u001b[0m\u001b[1mLayer (type)                   \u001b[0m\u001b[1m \u001b[0m┃\u001b[1m \u001b[0m\u001b[1mOutput Shape          \u001b[0m\u001b[1m \u001b[0m┃\u001b[1m \u001b[0m\u001b[1m      Param #\u001b[0m\u001b[1m \u001b[0m┃\n┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩\n│ mobilenetv2_1.00_192            │ (\u001b[38;5;45mNone\u001b[0m, \u001b[38;5;34m6\u001b[0m, \u001b[38;5;34m6\u001b[0m, \u001b[38;5;34m1280\u001b[0m)     │     \u001b[38;5;34m2,257,984\u001b[0m │\n│ (\u001b[38;5;33mFunctional\u001b[0m)                    │                        │               │\n├─────────────────────────────────┼────────────────────────┼───────────────┤\n│ global_average_pooling2d        │ (\u001b[38;5;45mNone\u001b[0m, \u001b[38;5;34m1280\u001b[0m)           │             \u001b[38;5;34m0\u001b[0m │\n│ (\u001b[38;5;33mGlobalAveragePooling2D\u001b[0m)        │                        │               │\n├─────────────────────────────────┼────────────────────────┼───────────────┤\n│ dense (\u001b[38;5;33mDense\u001b[0m)                   │ (\u001b[38;5;45mNone\u001b[0m, \u001b[38;5;34m104\u001b[0m)            │       \u001b[38;5;34m133,224\u001b[0m │\n└─────────────────────────────────┴────────────────────────┴───────────────┘\n","text/html":"<pre style=\"white-space:pre;overflow-x:auto;line-height:normal;font-family:Menlo,'DejaVu Sans Mono',consolas,'Courier New',monospace\">┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓\n┃<span style=\"font-weight: bold\"> Layer (type)                    </span>┃<span style=\"font-weight: bold\"> Output Shape           </span>┃<span style=\"font-weight: bold\">       Param # </span>┃\n┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩\n│ mobilenetv2_1.00_192            │ (<span style=\"color: #00d7ff; text-decoration-color: #00d7ff\">None</span>, <span style=\"color: #00af00; text-decoration-color: #00af00\">6</span>, <span style=\"color: #00af00; text-decoration-color: #00af00\">6</span>, <span style=\"color: #00af00; text-decoration-color: #00af00\">1280</span>)     │     <span style=\"color: #00af00; text-decoration-color: #00af00\">2,257,984</span> │\n│ (<span style=\"color: #0087ff; text-decoration-color: #0087ff\">Functional</span>)                    │                        │               │\n├─────────────────────────────────┼────────────────────────┼───────────────┤\n│ global_average_pooling2d        │ (<span style=\"color: #00d7ff; text-decoration-color: #00d7ff\">None</span>, <span style=\"color: #00af00; text-decoration-color: #00af00\">1280</span>)           │             <span style=\"color: #00af00; text-decoration-color: #00af00\">0</span> │\n│ (<span style=\"color: #0087ff; text-decoration-color: #0087ff\">GlobalAveragePooling2D</span>)        │                        │               │\n├─────────────────────────────────┼────────────────────────┼───────────────┤\n│ dense (<span style=\"color: #0087ff; text-decoration-color: #0087ff\">Dense</span>)                   │ (<span style=\"color: #00d7ff; text-decoration-color: #00d7ff\">None</span>, <span style=\"color: #00af00; text-decoration-color: #00af00\">104</span>)            │       <span style=\"color: #00af00; text-decoration-color: #00af00\">133,224</span> │\n└─────────────────────────────────┴────────────────────────┴───────────────┘\n</pre>\n"},"metadata":{}},{"output_type":"display_data","data":{"text/plain":"\u001b[1m Total params: \u001b[0m\u001b[38;5;34m2,391,208\u001b[0m (9.12 MB)\n","text/html":"<pre style=\"white-space:pre;overflow-x:auto;line-height:normal;font-family:Menlo,'DejaVu Sans Mono',consolas,'Courier New',monospace\"><span style=\"font-weight: bold\"> Total params: </span><span style=\"color: #00af00; text-decoration-color: #00af00\">2,391,208</span> (9.12 MB)\n</pre>\n"},"metadata":{}},{"output_type":"display_data","data":{"text/plain":"\u001b[1m Trainable params: \u001b[0m\u001b[38;5;34m133,224\u001b[0m (520.41 KB)\n","text/html":"<pre style=\"white-space:pre;overflow-x:auto;line-height:normal;font-family:Menlo,'DejaVu Sans Mono',consolas,'Courier New',monospace\"><span style=\"font-weight: bold\"> Trainable params: </span><span style=\"color: #00af00; text-decoration-color: #00af00\">133,224</span> (520.41 KB)\n</pre>\n"},"metadata":{}},{"output_type":"display_data","data":{"text/plain":"\u001b[1m Non-trainable params: \u001b[0m\u001b[38;5;34m2,257,984\u001b[0m (8.61 MB)\n","text/html":"<pre style=\"white-space:pre;overflow-x:auto;line-height:normal;font-family:Menlo,'DejaVu Sans Mono',consolas,'Courier New',monospace\"><span style=\"font-weight: bold\"> Non-trainable params: </span><span style=\"color: #00af00; text-decoration-color: #00af00\">2,257,984</span> (8.61 MB)\n</pre>\n"},"metadata":{}}],"execution_count":9},{"cell_type":"markdown","source":"## Train","metadata":{}},{"cell_type":"code","source":"t0 = time.time()\nmodel.fit(\n    get_training_dataset(),\n    steps_per_epoch=STEPS_PER_EPOCH,\n    epochs=EPOCHS,\n    validation_data=get_validation_dataset(),\n    verbose=2,\n)\ntrain_seconds = time.time() - t0\nprint(f\"train_seconds={train_seconds:.1f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:11:31.104122Z","iopub.execute_input":"2026-09-27T11:11:31.104481Z","iopub.status.idle":"2026-09-27T11:12:31.238788Z","shell.execute_reply.started":"2026-09-27T11:11:31.104455Z","shell.execute_reply":"2026-09-27T11:12:31.237979Z"}},"outputs":[{"name":"stdout","text":"Epoch 1/3\n","output_type":"stream"},{"name":"stderr","text":"WARNING: All log messages before absl::InitializeLog() is called are written to STDERR\nI0000 00:00:1790507497.662476     136 service.cc:152] XLA service 0x7d4dcc0e2bc0 initialized for platform CUDA (this does not guarantee that XLA will be used). Devices:\nI0000 00:00:1790507497.662516     136 service.cc:160]   StreamExecutor device (0): Tesla T4, Compute Capability 7.5\nI0000 00:00:1790507497.662522     136 service.cc:160]   StreamExecutor device (1): Tesla T4, Compute Capability 7.5\nI0000 00:00:1790507498.648417     136 cuda_dnn.cc:529] Loaded cuDNN version 91002\n2026-09-27 11:11:47.753433: E external/local_xla/xla/stream_executor/cuda/cuda_timer.cc:86] Delay kernel timed out: measured time has sub-optimal accuracy. There may be a missing warmup execution, please investigate in Nsight Systems.\n2026-09-27 11:11:47.900823: E external/local_xla/xla/stream_executor/cuda/cuda_timer.cc:86] Delay kernel timed out: measured time has sub-optimal accuracy. There may be a missing warmup execution, please investigate in Nsight Systems.\n2026-09-27 11:11:48.038106: E external/local_xla/xla/stream_executor/cuda/cuda_timer.cc:86] Delay kernel timed out: measured time has sub-optimal accuracy. There may be a missing warmup execution, please investigate in Nsight Systems.\nI0000 00:00:1790507509.841917     136 device_compiler.h:188] Compiled cluster using XLA!  This line is logged at most once for the lifetime of the process.\n","output_type":"stream"},{"name":"stdout","text":"398/398 - 35s - 89ms/step - accuracy: 0.5750 - loss: 1.8427 - val_accuracy: 0.7045 - val_loss: 1.1604\nEpoch 2/3\n","output_type":"stream"},{"name":"stderr","text":"/usr/local/lib/python3.12/dist-packages/keras/src/trainers/epoch_iterator.py:164: UserWarning: Your input ran out of data; interrupting training. Make sure that your dataset or generator can generate at least `steps_per_epoch * epochs` batches. You may need to use the `.repeat()` function when building your dataset.\n  self._interrupted_warning()\n","output_type":"stream"},{"name":"stdout","text":"398/398 - 12s - 30ms/step - accuracy: 0.7918 - loss: 0.8195 - val_accuracy: 0.7530 - val_loss: 0.9552\nEpoch 3/3\n398/398 - 12s - 30ms/step - accuracy: 0.8581 - loss: 0.5728 - val_accuracy: 0.7546 - val_loss: 0.9066\ntrain_seconds=60.1\n","output_type":"stream"}],"execution_count":10},{"cell_type":"markdown","source":"## Holdout accuracy","metadata":{}},{"cell_type":"code","source":"val_loss, val_acc = model.evaluate(get_validation_dataset(), verbose=0)\nprint(f\"holdout_accuracy={val_acc:.4f} val_loss={val_loss:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:12:31.239703Z","iopub.execute_input":"2026-09-27T11:12:31.239981Z","iopub.status.idle":"2026-09-27T11:12:33.723365Z","shell.execute_reply.started":"2026-09-27T11:12:31.239959Z","shell.execute_reply":"2026-09-27T11:12:33.722705Z"}},"outputs":[{"name":"stdout","text":"holdout_accuracy=0.7546 val_loss=0.9066\n","output_type":"stream"}],"execution_count":11},{"cell_type":"markdown","source":"## Macro F1 on holdout\n\nPetals to the Metal scores submissions on macro F1; we record both holdout accuracy AND macro F1 locally.","metadata":{}},{"cell_type":"code","source":"per_class_tp = collections.Counter()\nper_class_fp = collections.Counter()\nper_class_fn = collections.Counter()\n\nfor batch_images, batch_labels in get_validation_dataset():\n    preds = model.predict(batch_images, verbose=0).argmax(axis=1)\n    labels = batch_labels.numpy()\n    for y, p in zip(labels, preds):\n        if y == p:\n            per_class_tp[int(y)] += 1\n        else:\n            per_class_fp[int(p)] += 1\n            per_class_fn[int(y)] += 1\n\nf1s = []\nfor cls in range(NUM_CLASSES):\n    tp = per_class_tp[cls]\n    fp = per_class_fp[cls]\n    fn = per_class_fn[cls]\n    if tp + fp == 0 or tp + fn == 0:\n        f1s.append(0.0)\n        continue\n    precision = tp / (tp + fp)\n    recall = tp / (tp + fn)\n    if precision + recall == 0:\n        f1s.append(0.0)\n    else:\n        f1s.append(2 * precision * recall / (precision + recall))\n\nmacro_f1 = float(np.mean(f1s))\nprint(f\"macro_f1={macro_f1:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:12:33.724149Z","iopub.execute_input":"2026-09-27T11:12:33.724521Z","iopub.status.idle":"2026-09-27T11:12:51.305147Z","shell.execute_reply.started":"2026-09-27T11:12:33.724497Z","shell.execute_reply":"2026-09-27T11:12:51.304437Z"}},"outputs":[{"name":"stdout","text":"macro_f1=0.7421\n","output_type":"stream"}],"execution_count":12},{"cell_type":"markdown","source":"## Test-set inference + submission.csv\n\nOn Kaggle's kernel runner this writes `/kaggle/working/submission.csv` (the kernel push artifact).\nLocally we skip test-set inference unless the local slice contains a `test/` directory (it does not by default).","metadata":{}},{"cell_type":"code","source":"try:\n    test_files = tf.io.gfile.glob(_split_glob(\"test\"))\nexcept Exception:\n    test_files = []\n\nif test_files:\n    test_ds = get_test_dataset()\n    test_images_ds = test_ds.map(lambda image, idnum: image)\n    test_ids_ds = test_ds.map(lambda image, idnum: idnum).unbatch()\n    probs = model.predict(test_images_ds, verbose=0)\n    preds = probs.argmax(axis=1)\n    ids = [b.decode(\"utf-8\") for b in next(iter(test_ids_ds.batch(1_000_000))).numpy()]\n    out_path = \"submission.csv\"\n    if Path(\"/kaggle/working\").exists():\n        out_path = \"/kaggle/working/submission.csv\"\n    with open(out_path, \"w\", encoding=\"utf-8\") as f:\n        f.write(\"id,label\\n\")\n        for i, p in zip(ids, preds):\n            f.write(f\"{i},{int(p)}\\n\")\n    print(f\"submission written to {out_path} (n={len(ids)})\")\nelse:\n    print(\n        \"no test/ split available locally — skipping submission.csv (Kaggle kernel will produce it)\"\n    )","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-27T11:12:51.306207Z","iopub.execute_input":"2026-09-27T11:12:51.306602Z","iopub.status.idle":"2026-09-27T11:13:10.174435Z","shell.execute_reply.started":"2026-09-27T11:12:51.306575Z","shell.execute_reply":"2026-09-27T11:13:10.173649Z"}},"outputs":[{"name":"stderr","text":"2026-09-27 11:13:05.274794: E external/local_xla/xla/stream_executor/cuda/cuda_timer.cc:86] Delay kernel timed out: measured time has sub-optimal accuracy. There may be a missing warmup execution, please investigate in Nsight Systems.\n2026-09-27 11:13:05.424919: E external/local_xla/xla/stream_executor/cuda/cuda_timer.cc:86] Delay kernel timed out: measured time has sub-optimal accuracy. There may be a missing warmup execution, please investigate in Nsight Systems.\n2026-09-27 11:13:05.562027: E external/local_xla/xla/stream_executor/cuda/cuda_timer.cc:86] Delay kernel timed out: measured time has sub-optimal accuracy. There may be a missing warmup execution, please investigate in Nsight Systems.\n","output_type":"stream"},{"name":"stdout","text":"submission written to /kaggle/working/submission.csv (n=7382)\n","output_type":"stream"}],"execution_count":13},{"cell_type":"markdown","source":"## Done.\n\nLocal sanity metrics are surfaced as the last two stdout lines (`holdout_accuracy=` and `macro_f1=`),\nwhich `scripts/m10_t002_local_train.py` parses to record into `experiment_runs` via `kca run record-metric`.","metadata":{}}]}