This is a YOLO26s segmentation model for manga page layout analysis. It detects and segments three region types needed by manga OCR and translation pipelines:
Fonte do modelo
Descrição da fonte
This is a YOLO26s segmentation model for manga page layout analysis. It detects and segments three region types needed by manga OCR and translation pipelines:
The model is intended for manga document-understanding workflows where page regions must be located before OCR, reading-order reconstruction, translation, inpainting, or human review.
Fontes
1 fonteVerificado 18 de ago.
Artefatos de modelo
2 artefatosTrechos de fonte
2 trechosThis model is an Ultralytics-compatible YOLO26s instance segmentation model trained on Manga109-derived segmentation data. It predicts bounding boxes, class IDs, confidence scores, and pixel masks for manga page regions.
yolo26s-seg.yaml)yolo26s-seg.ptbest.ptThe labels are stored in the model checkpoint and match the expected YOLO dataset names mapping:
| Class ID | Label | Description |
|---|---|---|
| 0 | frame | Manga page panel/frame regions, including bordered or visually separated panels |
| 1 | text | Visible text regions, usually the regions passed to OCR or translation post-processing |
| 2 | balloon | Speech balloons, thought bubbles, narration bubbles, or similar text containers |
Notes:
text is the visual text region, not the OCR transcription.balloon is the container region around dialogue or narration text.frame is the panel/layout region, not necessarily a semantic scene label.data.yaml, inference code, and downstream post-processing.Recommended dataset config:
names:
0: frame
1: text
2: balloon
Use this model to segment manga page regions from a page image. Direct outputs can be used to locate:
This model is designed to be one component in a larger manga translation or document-understanding system. A typical downstream flow is:
Example structured output expected by a downstream pipeline:
{
"page": "example_page.jpg",
"regions": [
{
"id": 1,
"class_id": 0,
"label": "frame",
"confidence": 0.94,
"bbox": [x1, y1, x2, y2],
"mask": "..."
},
{
"id": 2,
"class_id": 2,
"label": "balloon",
"confidence": 0.91,
"bbox": [x1, y1, x2, y2],
"mask": "..."
},
{
"id": 3,
"class_id": 1,
"label": "text",
"confidence": 0.89,
"bbox": [x1, y1, x2, y2],
"mask": "..."
}
]
}
This model should not be treated as:
It only segments visible page regions. It does not understand text content, speaker identity, story context, or translation quality.
Install dependencies:
pip install ultralytics pillow opencv-python
Run inference:
from ultralytics import YOLO
# Replace with your local path or the Hugging Face model ID after upload.
model = YOLO("best.pt")
results = model.predict(
source="example_manga_page.jpg",
imgsz=1280,
conf=0.25,
iou=0.7,
retina_masks=True,
)
class_names = {
0: "frame",
1: "text",
2: "balloon",
}
for result in results:
if result.boxes is None:
print("No regions detected.")
continue
for i, box in enumerate(result.boxes):
class_id = int(box.cls[0])
confidence = float(box.conf[0])
bbox = box.xyxy[0].tolist()
label = class_names.get(class_id, str(class_id))
print({
"index": i,
"class_id": class_id,
"label": label,
"confidence": confidence,
"bbox": bbox,
})
# Saves an annotated image with boxes/masks.
result.save(filename="segmented_output.jpg")
For high-quality mask extraction in a manga translation pipeline, use retina_masks=True during inference so masks are returned at higher resolution.
This model uses a merged Manga109-derived segmentation dataset with three region classes: frame, text, and balloon.
| Dataset | Hugging Face ID | Use | Notes |
|---|---|---|---|
| MangaSegmentation | MS92/MangaSegmentation | Segmentation annotations for manga regions | Dataset card references “Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 Dataset.” |
| Manga109 Region-Level Text Segmentation | ShadowB/Manga109_RegionLevelTextSegmentation | Region-level text masks | Used to support the text class and downstream OCR/translation needs. |
The provided split audit records a book-level split across 109 manga groups:
| Split | Books / Groups | Images |
|---|---|---|
| Train | 83 | 7,174 |
| Validation | 12 | 1,468 |
| Test | 14 | 1,488 |
| Total | 109 | 10,130 |
The book-level split is important because random page-level splits can overestimate performance by leaking manga-specific style, art, and layout patterns between train and validation data.
The training data was normalized into a YOLO-compatible segmentation layout with the following class mapping:
0: frame
1: text
2: balloon
Known preprocessing goals:
yolo segment train.The model was trained on Kaggle with Ultralytics YOLO26s segmentation. The training script builds a book-level train/validation/test split, maps the labels to three classes (frame, text, balloon), and keeps overlap_mask=False because manga regions can sit inside each other.
Training used yolo26s-seg.pt as the starting checkpoint, image size 1280, batch size 8 across two GPUs, and MuSGD. The run completed 41 epochs and took 11h 13m overall. The checkpoint stores an Ultralytics training time value of 11.0054 hours, which reflects the active training budget rather than the full notebook runtime.
| Hyperparameter | Value |
|---|---|
| Architecture | YOLO26s segmentation |
| Base checkpoint | yolo26s-seg.pt |
| Image size | 1280 |
| Batch size | 8 |
| Epochs completed | 41 |
| Overall run time | 11h 13m |
| Optimizer | MuSGD |
| Learning rate | 0.01 initial, 0.01 final factor |
| Momentum | 0.937 |
| Weight decay | 0.0005 |
| Warmup epochs | 3.0 |
| Cosine LR | True |
| AMP |
| Item | Value |
|---|---|
Checkpoint size, best.pt | 23,439,133 bytes |
| Parameters | 11,436,269 |
The available metrics are from the validation run recorded in best.pt, results.csv, and the validation artifacts in this repository.
/validationResultsOfMangaModel/results.csv and checkpoint train_metricsRecommended factors for further evaluation:
The following values come from the best.pt checkpoint train_metrics field:
| Metric | Value |
|---|---|
| Box Precision | 0.96521 |
| Box Recall | 0.95165 |
| Box mAP@0.5 | 0.97494 |
| Box mAP@0.5:0.95 | 0.89988 |
| Mask Precision | 0.96564 |
| Mask Recall | 0.95026 |
| Mask mAP@0.5 | 0.97013 |
| Mask mAP@0.5:0.95 | 0.84573 |
| Validation box loss | 0.43638 |
| Validation segmentation loss | 0.59429 |
| Validation classification loss | 0.26392 |
| Validation DFL loss | 0.00241 |
| Fitness |
The final row in results.csv, epoch 41, records very similar overall metrics:
| Metric | Epoch 41 Value |
|---|---|
| Box Precision | 0.96489 |
| Box Recall | 0.95021 |
| Box mAP@0.5 | 0.97432 |
| Box mAP@0.5:0.95 | 0.89907 |
| Mask Precision | 0.96627 |
| Mask Recall | 0.94811 |
| Mask mAP@0.5 | 0.96986 |
| Mask mAP@0.5:0.95 | 0.84459 |
The local artifacts provided here include overall metrics, PR/F1/P/R curves, labels visualization, and confusion matrices. A per-class numeric mAP table was not present in results.csv.
To add per-class metrics, run a validation command that prints or exports per-class results, then update this table:
| Class | Box mAP@0.5 | Box mAP@0.5:0.95 | Mask mAP@0.5 | Mask mAP@0.5:0.95 |
|---|---|---|---|---|
frame | TODO | TODO | TODO | TODO |
text | TODO | TODO | TODO | TODO |
balloon | TODO | TODO | TODO | TODO |
Suggested command when the dataset is available:
yolo segment val \
model=best.pt \
data=/path/to/data.yaml \
imgsz=1280 \...
--- license: mit github: https://github.com/sadowb/CuratorML datasets: - MS92/MangaSegmentation - ShadowB/Manga109_RegionLevelTextSegmentation language: - ja pipeline_tag: image-segmentation library_name: ultralytics tags: - manga - manga-segmentation - image-segmentation - instance-segmentation - yolo - yolo26 - yolo26s - comics - document-understanding - layout-analysis - ocr-preprocessing - translation-pipeline model-index: - name: yolo26s-manga-panel-bubble-text-segmentation results: - task: type: image-segmentation name: Manga region instance segmentation dataset: type: custom name: Book-level validation split from Manga109-derived segmentation data metrics: - type: precision name: Box Precision value: 0.96521 - type: recall name: Box Recall value: 0.95165 - type: map name: Box mAP@0.5 value: 0.97494 - type: map name: Box mAP@0.5:0.95 value: 0.89988 - type: precision name: Mask Precision value: 0.96564 - type: recall name: Mask Recall value: 0.95026 - type: map name: Mask mAP@0.5 value: 0.97013 - type: map name: Mask mAP@0.5:0.95 value: 0.84573 --- # YOLO26s Manga Panel, Text, and Balloon Segmentation This is a YOLO26s segmentation model for manga page layout analysis. It detects and segments three region types needed by manga OCR and translation pipelines: - panels / frames, - text regions, - speech or narration balloons. The model is intended for manga document-understanding workflows where page regions must be located before OCR, reading-order reconstruction, translation, inpainting, or human review. ## Model Details ### Model Description This model is an Ultralytics-compatible YOLO26s instance segmentation model trained on Manga109-derived segmentation data. It predicts bounding boxes, class IDs, confidence scores, and pixel masks for manga page regions. - **Developed by:** ShadowB / Abdelhadi Marjane - **Model type:** Image segmentation / instance segmentation - **Architecture:** YOLO26s segmentation model (`yolo26s-seg.yaml`) - **Base checkpoint:** `yolo26s-seg.pt` - **Library:** Ultralytics - **Task:** Manga region instance segmentation - **Primary domain:** Manga/comic page images - **Languages:** Japanese manga pages. The model detects page regions visually; it does not read or translate text. - **License:** MIT for this model repository. Dataset licenses and access rules may differ. - **Number of clas...
Source context: 1039 downloads · 7 likes · Pipeline image-segmentation · Library ultralytics · Repo ShadowB/Manga109-panel-balloon-text-yolov26-segmentation
| True |
| Device | 0,1 |
| Overlap mask | False |
| Main augmentations | mosaic 0.3, copy-paste 0.1, HSV-V 0.04, no flips/rotation |
| 1.74561 |