{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceType":"competition","sourceId":132872,"databundleVersionId":15931798,"isSourceIdPinned":false}],"dockerImageVersionId":31287,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport torch\nimport math\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport os\nfrom pathlib import Path\nimport zipfile\nfrom tqdm import tqdm\nfrom datetime import datetime, timedelta\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:02:57.657374Z","iopub.execute_input":"2026-03-05T13:02:57.657665Z","iopub.status.idle":"2026-03-05T13:02:57.662061Z","shell.execute_reply.started":"2026-03-05T13:02:57.657641Z","shell.execute_reply":"2026-03-05T13:02:57.661372Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Вам надо сделать вот эти пункты:\n1) Зафиксировать сиды\n2) Сделать валидацию\n3) Написано усреднение моделей\n\nВ здаче вам даны векторы  X : N_L, H, W, T  (N_L - 5 векторов фичей (разные длины волн)) размерность T (t-2, t-1, t, t+1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T11:10:33.509320Z","iopub.execute_input":"2026-03-05T11:10:33.509650Z","iopub.status.idle":"2026-03-05T11:10:33.516223Z","shell.execute_reply.started":"2026-03-05T11:10:33.509626Z","shell.execute_reply":"2026-03-05T11:10:33.514205Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Интерпретация постановки задачи\n\nПеред нами задача **бинарной сегментации спутниковых изображений**.  \nНеобходимо предсказать **маску осадков** — определить для каждого пикселя изображения, наблюдается ли в этой точке дождь.\n\nМодель получает многоканальное изображение и должна предсказать бинарную маску:\n\n\nimage → CNN → mask\n\n\nгде каждому пикселю соответствует значение:\n\n\n0 — нет осадков\n1 — дождь\n\n\nТаким образом задача формулируется как **pixel-wise binary classification**.\n\n---\n\n# Формат данных\n\nИсходные данные имеют структуру:\n\n\nX : N_L, H, W, T\n\n\nгде\n\n\nN_L — число длин волн (спектральные каналы)\nT — временные наблюдения\n\n\nВ задаче:\n\n\nN_L = 5\nT = 4\n\n\nЭто означает, что каждый пример содержит наблюдения **5 спектральных каналов в 4 момента времени**.\n\nИсходный формат данных:\n\n\n[N, 5, H, W, 4]\n\n\nПосле преобразования данные можно рассматривать как **многоканальное изображение**.\n\nБазовое представление:\n\n\n5 × 4 = 20 каналов\n\n\nДополнительно были построены **временные разности между соседними наблюдениями**, что позволило получить дополнительные признаки динамики облаков.\n\nИтоговое число каналов:\n\n\n20 исходных + 15 временных разностей = 35 каналов\n\n\nТаким образом вход модели имеет вид:\n\n\n[35, 256, 256]\n\n\n---\n\n# Целевая переменная\n\nЦелевая переменная — **бинарная маска осадков**.\n\nФорма:\n\n\nY : [H, W]\n\n\nПри обучении:\n\n\n[1, H, W]\n\n\n---\n\n# Метрика\n\nКачество решения оценивается метрикой **Dice coefficient**.\n\nФормула:\n\n\nDice = 2TP / (2TP + FP + FN)\n\n\nМетрика измеряет степень совпадения между предсказанной и истинной масками.\n\n---\n\n# Базовое решение\n\nBaseline использует сегментационную архитектуру **U-Net**, обучаемую на подготовленных изображениях.\n\nОднако baseline содержит упрощения и служит отправной точкой для улучшений.\n\n---\n\n# Что было добавлено в решении\n\nВ ходе работы были внесены следующие изменения относительно baseline:\n\n**1. Использование всей информации входных данных**\n\nВместо ручного формирования RGB-подобных изображений используются **все доступные каналы данных**, а также дополнительные признаки временной динамики.\n\nИтоговый вход модели:\n\n\n35 каналов\n\n\n---\n\n**2. Архитектура модели**\n\nИспользуется современная реализация сегментационной модели:\n\n\nUNet++\n\n\nЭто более мощная модификация U-Net, позволяющая лучше использовать многоуровневые признаки.\n\n---\n\n**3. Аугментации данных**\n\nДля повышения устойчивости модели используются базовые аугментации:\n\n- горизонтальное отражение\n- вертикальное отражение\n- повороты\n\n---\n\n**4. Валидационное разбиение**\n\nДанные разделяются на обучающую и валидационную выборки, что позволяет контролировать качество модели в процессе обучения.\n\n---\n\n**5. Подбор оптимального порога**\n\nМодель предсказывает **вероятности принадлежности пикселя классу \"дождь\"**.\n\nДля получения бинарной маски выполняется подбор оптимального порога вероятности на валидационной выборке.\n\n---\n\n# Итоговый пайплайн\n\nИтоговый процесс решения задачи выглядит следующим образом:\n\n1. Подготовка многоканального входного изображения  \n2. Формирование дополнительных временных признаков  \n3. Обучение сегментационной модели UNet++  \n4. Получение вероятностных карт  \n5. Подбор оптимального порога бинаризации  \n6. Формирование финальных масок осадков","metadata":{}},{"cell_type":"markdown","source":"# Коротко:\nУбираем handcraft, выполняем указанные шаги и принимаем всю информацию:\n\n## Сейчас baseline делает:\n\n20 каналов\n   ↓\nhandcrafted features\n   ↓\n3 канала\n   ↓\nUNet\n\n## Но можно сделать намного лучше:\n\n20 каналов\n   ↓\nUNet(in_channels=16)","metadata":{}},{"cell_type":"code","source":"import albumentations as A\n\ntrain_transform = A.Compose([\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=0.5),\n    \n    A.Affine(\n        scale=(0.9, 1.1),\n        translate_percent=(0.0, 0.05),\n        rotate=(-20, 20),\n        p=0.5\n    ),\n])\n\nval_transform = A.Compose([])  # можно и None, но так проще и единообразно","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:03.370383Z","iopub.execute_input":"2026-03-05T13:03:03.370673Z","iopub.status.idle":"2026-03-05T13:03:03.377028Z","shell.execute_reply.started":"2026-03-05T13:03:03.370650Z","shell.execute_reply":"2026-03-05T13:03:03.376385Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"val_transform = A.Compose([])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:03.887235Z","iopub.execute_input":"2026-03-05T13:03:03.887521Z","iopub.status.idle":"2026-03-05T13:03:03.891515Z","shell.execute_reply.started":"2026-03-05T13:03:03.887495Z","shell.execute_reply":"2026-03-05T13:03:03.890718Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Нормализация и вычитание потоков (не для трех как в бейзе, а чуть иначе)","metadata":{}},{"cell_type":"code","source":"class SDataset(Dataset):\n\n    def __init__(self, X, Y=None, transform=None):\n        self.X = X\n        self.Y = Y\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.X)\n\n    def __getitem__(self, idx):\n\n        x = self.X[idx].cpu().numpy().astype(\"float32\")\n        x = np.transpose(x, (1,2,0,3))\n\n        H, W, L, T = x.shape\n\n        diffs = []\n        for l in range(L):\n            for t in range(T-1):\n                diffs.append(x[:,:,l,t+1] - x[:,:,l,t])\n\n        diffs = np.stack(diffs, axis=-1)\n\n        x = np.concatenate(\n            [\n                x.reshape(H,W,L*T),\n                diffs\n            ],\n            axis=-1\n        )\n\n        x = x / 256.0\n\n        if self.Y is None:\n\n            if self.transform:\n                x = self.transform(image=x)[\"image\"]\n\n            x = torch.from_numpy(x).permute(2,0,1).float()\n            return x\n\n        y = self.Y[idx].cpu().numpy().astype(\"float32\")\n        y = np.squeeze(y)\n\n        if self.transform:\n            aug = self.transform(image=x, mask=y)\n            x = aug[\"image\"]\n            y = aug[\"mask\"]\n\n        y = (y > 0.5).astype(\"float32\")\n\n        x = torch.from_numpy(x).permute(2,0,1).float()\n        y = torch.from_numpy(y).unsqueeze(0).float()\n\n        return x, y","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:27.417044Z","iopub.execute_input":"2026-03-05T13:03:27.417340Z","iopub.status.idle":"2026-03-05T13:03:27.425633Z","shell.execute_reply.started":"2026-03-05T13:03:27.417313Z","shell.execute_reply":"2026-03-05T13:03:27.425015Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_all = torch.from_numpy(np.load('/kaggle/input/competitions/vsosh-pesok-cv/X_train.npy'))\nY_all = torch.from_numpy(np.load('/kaggle/input/competitions/vsosh-pesok-cv/y_train.npy'))\nX_test_raw = torch.from_numpy(np.load('/kaggle/input/competitions/vsosh-pesok-cv/X_test.npy'))\n\nprint(\"X_all:\", X_all.shape)      # [N, 5, 256, 256, 4]\nprint(\"Y_all:\", Y_all.shape)      # ожидаем [N, 256, 256] (или [N,256,256,1] — обработаем ниже)\nprint(\"X_test:\", X_test_raw.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:33.781773Z","iopub.execute_input":"2026-03-05T13:03:33.782098Z","iopub.status.idle":"2026-03-05T13:03:35.239472Z","shell.execute_reply.started":"2026-03-05T13:03:33.782075Z","shell.execute_reply":"2026-03-05T13:03:35.238526Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if Y_all.ndim == 4 and Y_all.shape[-1] == 1:\n    Y_all = Y_all[..., 0]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:37.326568Z","iopub.execute_input":"2026-03-05T13:03:37.327341Z","iopub.status.idle":"2026-03-05T13:03:37.330860Z","shell.execute_reply.started":"2026-03-05T13:03:37.327312Z","shell.execute_reply":"2026-03-05T13:03:37.330162Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Добавим возможность стратификации по факту наличия дождя для стабильности обучения","metadata":{}},{"cell_type":"code","source":"rain_label = (Y_all.sum(dim=(1,2)) > 0).numpy()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:38.251616Z","iopub.execute_input":"2026-03-05T13:03:38.251936Z","iopub.status.idle":"2026-03-05T13:03:38.491074Z","shell.execute_reply.started":"2026-03-05T13:03:38.251911Z","shell.execute_reply":"2026-03-05T13:03:38.490435Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nX_train_raw, X_val_raw, y_train, y_val = train_test_split(\n    X_all, Y_all,\n    test_size=0.2,\n    random_state=42,\n    stratify=rain_label\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:39.726726Z","iopub.execute_input":"2026-03-05T13:03:39.727357Z","iopub.status.idle":"2026-03-05T13:03:41.041939Z","shell.execute_reply.started":"2026-03-05T13:03:39.727328Z","shell.execute_reply":"2026-03-05T13:03:41.041118Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"y min/max:\", y_train.min().item(), y_train.max().item())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:41.211558Z","iopub.execute_input":"2026-03-05T13:03:41.212167Z","iopub.status.idle":"2026-03-05T13:03:41.226063Z","shell.execute_reply.started":"2026-03-05T13:03:41.212131Z","shell.execute_reply":"2026-03-05T13:03:41.225318Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"trainl = DataLoader(SDataset(X_train_raw, y_train, transform=train_transform), batch_size=16, shuffle=True)\nvall   = DataLoader(SDataset(X_val_raw,   y_val,   transform=val_transform),   batch_size=16, shuffle=False)\n\ntestl  = DataLoader(SDataset(X_test_raw,  Y=None,  transform=val_transform),   batch_size=16, shuffle=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:41.904774Z","iopub.execute_input":"2026-03-05T13:03:41.905100Z","iopub.status.idle":"2026-03-05T13:03:42.071969Z","shell.execute_reply.started":"2026-03-05T13:03:41.905075Z","shell.execute_reply":"2026-03-05T13:03:42.071370Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pip install segmentation-models-pytorch","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:44.592723Z","iopub.execute_input":"2026-03-05T13:03:44.593040Z","iopub.status.idle":"2026-03-05T13:03:48.048176Z","shell.execute_reply.started":"2026-03-05T13:03:44.593014Z","shell.execute_reply":"2026-03-05T13:03:48.047159Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import segmentation_models_pytorch as smp","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:48.049968Z","iopub.execute_input":"2026-03-05T13:03:48.050308Z","iopub.status.idle":"2026-03-05T13:03:48.054047Z","shell.execute_reply.started":"2026-03-05T13:03:48.050282Z","shell.execute_reply":"2026-03-05T13:03:48.053264Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = smp.UnetPlusPlus(\n    encoder_name=\"resnet34\",\n    encoder_weights=\"imagenet\",\n    in_channels=35,\n    classes=1\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:48.055048Z","iopub.execute_input":"2026-03-05T13:03:48.055321Z","iopub.status.idle":"2026-03-05T13:03:48.675423Z","shell.execute_reply.started":"2026-03-05T13:03:48.055288Z","shell.execute_reply":"2026-03-05T13:03:48.674626Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def dice_loss(logits, target, eps=1e-6):\n\n    prob = torch.sigmoid(logits)\n\n    intersection = (prob * target).sum(dim=(1,2,3))\n    union = prob.sum(dim=(1,2,3)) + target.sum(dim=(1,2,3))\n\n    dice = (2 * intersection + eps) / (union + eps)\n\n    return 1 - dice.mean()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:48.677008Z","iopub.execute_input":"2026-03-05T13:03:48.677410Z","iopub.status.idle":"2026-03-05T13:03:48.681727Z","shell.execute_reply.started":"2026-03-05T13:03:48.677378Z","shell.execute_reply":"2026-03-05T13:03:48.681034Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"bce = nn.BCEWithLogitsLoss()\n\ndef loss_fn(logits, y):\n    return bce(logits, y) + dice_loss(logits, y)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:48.682582Z","iopub.execute_input":"2026-03-05T13:03:48.682943Z","iopub.status.idle":"2026-03-05T13:03:48.694127Z","shell.execute_reply.started":"2026-03-05T13:03:48.682922Z","shell.execute_reply":"2026-03-05T13:03:48.693498Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"device = 'cuda'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:50.020627Z","iopub.execute_input":"2026-03-05T13:03:50.021005Z","iopub.status.idle":"2026-03-05T13:03:50.024400Z","shell.execute_reply.started":"2026-03-05T13:03:50.020977Z","shell.execute_reply":"2026-03-05T13:03:50.023711Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x,y = next(iter(trainl))\n\nprint(\"x shape\",x.shape)\nprint(\"y unique\",np.unique(y.numpy()))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:03:50.294770Z","iopub.execute_input":"2026-03-05T13:03:50.295316Z","iopub.status.idle":"2026-03-05T13:03:50.911583Z","shell.execute_reply.started":"2026-03-05T13:03:50.295286Z","shell.execute_reply":"2026-03-05T13:03:50.910745Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model = model.to(device)\n\noptimizer = torch.optim.Adam(model.parameters(), lr=1e-3)\n\nEPOCHS = 15\n\nfor epoch in range(EPOCHS):\n\n    model.train()\n\n    train_loss = 0\n\n    for x,y in trainl:\n\n        x = x.to(device)\n        y = y.to(device)\n\n        optimizer.zero_grad()\n\n        logits = model(x)\n\n        loss = loss_fn(logits, y)\n\n        loss.backward()\n\n        optimizer.step()\n\n        train_loss += loss.item()\n\n    print(f\"epoch {epoch} train_loss {train_loss/len(trainl):.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:19:48.354701Z","iopub.execute_input":"2026-03-05T13:19:48.355315Z","iopub.status.idle":"2026-03-05T13:30:18.795765Z","shell.execute_reply.started":"2026-03-05T13:19:48.355284Z","shell.execute_reply":"2026-03-05T13:30:18.795120Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def dice_score(logits, target, eps=1e-6):\n\n    pred = torch.sigmoid(logits)\n    pred = (pred > 0.5).float()\n\n    intersection = (pred * target).sum(dim=(1,2,3))\n    union = pred.sum(dim=(1,2,3)) + target.sum(dim=(1,2,3))\n\n    dice = (2*intersection + eps) / (union + eps)\n\n    return dice.mean().item()\n\ndef dice_score_soft(logits, target, eps=1e-6):\n    prob = torch.sigmoid(logits)\n    inter = (prob * target).sum(dim=(1,2,3))\n    union = prob.sum(dim=(1,2,3)) + target.sum(dim=(1,2,3))\n    dice = (2*inter + eps) / (union + eps)\n    return dice.mean().item()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:04:34.043255Z","iopub.execute_input":"2026-03-05T13:04:34.043559Z","iopub.status.idle":"2026-03-05T13:04:34.049192Z","shell.execute_reply.started":"2026-03-05T13:04:34.043536Z","shell.execute_reply":"2026-03-05T13:04:34.048451Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"with torch.no_grad():\n    x, y = next(iter(vall))\n    x = x.to(device)\n    probs = torch.sigmoid(model(x))\n\nprint(\"probs min/max/mean:\", probs.min().item(), probs.max().item(), probs.mean().item())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:30:51.335061Z","iopub.execute_input":"2026-03-05T13:30:51.335657Z","iopub.status.idle":"2026-03-05T13:30:51.779501Z","shell.execute_reply.started":"2026-03-05T13:30:51.335628Z","shell.execute_reply":"2026-03-05T13:30:51.778689Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.eval()\nsoft_total = 0\nhard_total = 0\nn = 0\n\nwith torch.no_grad():\n    for x,y in vall:\n        x=x.to(device); y=y.to(device)\n        logits = model(x)\n        soft_total += dice_score_soft(logits,y)\n        hard_total += dice_score(logits,y)  # твоя (threshold=0.5)\n        n += 1\n\nprint(\"VAL soft dice:\", soft_total/n)\nprint(\"VAL hard dice@0.5:\", hard_total/n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:30:55.502474Z","iopub.execute_input":"2026-03-05T13:30:55.502761Z","iopub.status.idle":"2026-03-05T13:31:01.032504Z","shell.execute_reply.started":"2026-03-05T13:30:55.502736Z","shell.execute_reply":"2026-03-05T13:31:01.031866Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"fraction >0.5:\", (probs > 0.5).mean().item())\nprint(\"fraction >0.2:\", (probs > 0.2).mean().item())\nprint(\"fraction >0.1:\", (probs > 0.1).mean().item())","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.eval()\n\npred_list = []\n\nwith torch.no_grad():\n\n    for x in testl:\n\n        x = x.to(device)\n\n        logits = model(x)\n\n        probs = torch.sigmoid(logits).cpu().numpy()\n\n        pred_list.extend(probs)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Подбираем оптимальный порог","metadata":{}},{"cell_type":"code","source":"def find_best_threshold(model, loader):\n\n    thresholds = np.arange(0.01, 0.31, 0.05)\n\n    best_t = 0.5\n    best_dice = 0\n\n    model.eval()\n\n    with torch.no_grad():\n\n        for t in thresholds:\n\n            dices = []\n\n            for x,y in loader:\n\n                x = x.to(device)\n                y = y.to(device)\n\n                logits = model(x)\n\n                pred = (torch.sigmoid(logits) > t).float()\n\n                intersection = (pred*y).sum(dim=(1,2,3))\n                union = pred.sum(dim=(1,2,3)) + y.sum(dim=(1,2,3))\n\n                dice = (2*intersection/(union+1e-6)).mean()\n\n                dices.append(dice.item())\n\n            d = np.mean(dices)\n\n            if d > best_dice:\n                best_dice = d\n                best_t = t\n\n    print(\"best threshold:\",best_t,\"dice:\",best_dice)\n\n    return best_t","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:31:20.993083Z","iopub.execute_input":"2026-03-05T13:31:20.993709Z","iopub.status.idle":"2026-03-05T13:31:20.999765Z","shell.execute_reply.started":"2026-03-05T13:31:20.993682Z","shell.execute_reply":"2026-03-05T13:31:20.998982Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"best_t = find_best_threshold(model,vall)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:31:21.694647Z","iopub.execute_input":"2026-03-05T13:31:21.694958Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"masks = []\n\nfor p in pred_list:\n\n    m = (p[0] > best_t).astype(np.uint8)\n\n    masks.append(m)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:32:46.492641Z","iopub.execute_input":"2026-03-05T13:32:46.493397Z","iopub.status.idle":"2026-03-05T13:32:46.542544Z","shell.execute_reply.started":"2026-03-05T13:32:46.493369Z","shell.execute_reply":"2026-03-05T13:32:46.541996Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Остатки бейзлайна (которые используются в решении):","metadata":{}},{"cell_type":"code","source":"def rle_encode(x, fg_val=1):\n    \"\"\"\n    Args:\n        x:  numpy array of shape (height, width), 1 - mask, 0 - background\n    Returns: run length encoding as list\n    \"\"\"\n\n    dots = np.where(\n        x.T.flatten() == fg_val)[0]  # .T sets Fortran order down-then-right\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if b > prev + 1:\n            run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n    return run_lengths\n\n\ndef list_to_string(x):\n    \"\"\"\n    Converts list to a string representation\n    Empty list returns '-'\n    \"\"\"\n    if not x:\n        return \"-\"\n    return \" \".join(map(str, x))\n\n\ndef rle_decode(mask_rle, shape=(256, 256)):\n    '''\n    mask_rle: run-length as string formatted (start length)\n              empty predictions need to be encoded with '-'\n    shape: (height, width) of array to return \n    Returns numpy array, 1 - mask, 0 - background\n    '''\n\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    if mask_rle != '-': \n        s = mask_rle.split()\n        starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n        starts -= 1\n        ends = starts + lengths\n        for lo, hi in zip(starts, ends):\n            img[lo:hi] = 1\n    return img.reshape(shape, order='F') ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:04:40.087726Z","iopub.status.idle":"2026-03-05T13:04:40.088071Z","shell.execute_reply.started":"2026-03-05T13:04:40.087903Z","shell.execute_reply":"2026-03-05T13:04:40.087926Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.DataFrame()\n\nsubmission[\"id\"] = np.arange(len(masks))\nsubmission[\"encoded_pixels\"] = \"\"\n\nfor i, mask in enumerate(masks):\n\n    submission.loc[i, \"encoded_pixels\"] = list_to_string(\n        rle_encode(mask)\n    )\n\nsubmission.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:32:50.722662Z","iopub.execute_input":"2026-03-05T13:32:50.723278Z","iopub.status.idle":"2026-03-05T13:33:00.769616Z","shell.execute_reply.started":"2026-03-05T13:32:50.723249Z","shell.execute_reply":"2026-03-05T13:33:00.768903Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"assert len(submission) == len(X_test_raw), (len(submission), len(X_test_raw))\nassert submission[\"id\"].is_unique\nassert submission[\"id\"].min() == 0\nassert submission[\"id\"].max() == len(X_test_raw) - 1","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:16:04.187039Z","iopub.execute_input":"2026-03-05T13:16:04.187681Z","iopub.status.idle":"2026-03-05T13:16:04.192198Z","shell.execute_reply.started":"2026-03-05T13:16:04.187652Z","shell.execute_reply":"2026-03-05T13:16:04.191480Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"empty_rate = (submission[\"encoded_pixels\"] == \"-\").mean()\nprint(\"empty_rate:\", empty_rate)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:33:18.645767Z","iopub.execute_input":"2026-03-05T13:33:18.646091Z","iopub.status.idle":"2026-03-05T13:33:18.650751Z","shell.execute_reply.started":"2026-03-05T13:33:18.646065Z","shell.execute_reply":"2026-03-05T13:33:18.650179Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission.to_csv('sub5.csv', index = None)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-03-05T13:33:22.486182Z","iopub.execute_input":"2026-03-05T13:33:22.486906Z","iopub.status.idle":"2026-03-05T13:33:22.491902Z","shell.execute_reply.started":"2026-03-05T13:33:22.486875Z","shell.execute_reply":"2026-03-05T13:33:22.491203Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}